Semantic Experiment Graph
Make Every Marketing Test Teach the Next One
After Reading This Ebook, You Will:
- ✓ Separate winner-picking from mechanism learning
- ✓ Preregister semantic and surface compositions and file multi-type receipts
- ✓ Design controlled descendants that isolate carrier, concept, or structure
- ✓ Walk heat to graph-guided next tests and promote a marketing canon without letting likes rewrite truth
TL;DR
- • A/B testing crowns a winner; a semantic experiment graph designs the next race to find out why.
- • Claim-grain stimuli make response addressable at concept level.
- • File two compositions + multi-type receipts; treat heat as a multi-dimensional nudge, never as truth.
- • Isolate carrier / concept / structure; walk convergence for hybrids; promote a marketing canon through a staged ladder.
- • The winning creative is perishable. The experiment graph compounds.
Winner-Picking Is Not Learning
Most teams that “do experimentation” are holding a tournament. The dashboard crowns a variant. The organisation clones its surface. Almost nothing compounds as a mechanism.
Picture the Monday meeting. Two subject lines, two creatives, two landing-page layouts. Traffic split cleanly enough. Variant B beat variant A on the metric the dashboard was built to optimise. Someone screenshots the green uplift. Someone else opens a brief that says: make more like B. By Wednesday the calendar is full of surface cousins — same colour family, same celebrity doorway, same length band — and the organisation calls that a culture of testing.
It is a culture of selection. Selection is not nothing. Randomised assignment still earns its keep: it controls seasonality, traffic mix, and the other gremlins that make before/after stories dishonest.1 Industry tooling is built around that story — compare variants, pick a better performer, ship it.2 The problem is what the organisation stores after the tournament ends. It stores a winner label. It almost never stores a mechanism.
That is why the next test is usually a clone. Clones are what you invent when the only thing you learned is “this object scored.” You do not know which idea inside the object did the work. You do not know which rhetorical move carried it. You do not know whether the image helped, hurt, or was mostly noise. You especially do not know what neighbouring idea might inherit the effect if you stopped worshipping the surface.
The unit of analysis is wrong
A pillar post or a multi-block ad is a confounded instrument. Concept, vehicle, structure, rendering, channel, timing, and audience all collapse into one score. The dashboard reports that “the post did well.” Nothing attributes. Averages destroy edges on the measurement side the same way they destroy edges on the meaning side. You walk away with folklore: black-and-white images work; Einstein thumbnails perform; Tuesdays convert; long posts beat short posts — until they do not.
Some of those stories may even be directionally useful in a narrow window. None of them is a theory you can compound. None of them tells a new hire, six months later, what to isolate next.
User-research literature has been saying a cousin of this for years from another discipline: A/B methods can show behavioural difference without explaining the qualitative why. You can watch time-on-page move and still not know what users believed, feared, or recognised.3 Marketers often treat that as a reason to run “more qualitative research someday.” The sharper move is architectural: stop forcing every experiment to measure multi-idea artefacts as if they were atomic.
What generative systems made worse — and better
Variant production used to be expensive enough that teams could pretend scarcity of creatives was the bottleneck. Generative systems deleted that pretence. You can now mint a hundred near-misses before lunch. That should have made learning explode. In practice it often makes mutation theatre explode: more objects, same ignorance, prettier decks.
The scarce resource moved. It is no longer the ability to produce another candidate. It is the ability to design a candidate that would falsify a mechanism rather than redecorate a winner. That is a harder skill. It is also a skill that compounds if you instrument it.
A/B testing crowns a winner. A semantic experiment graph teaches the compiler why the race may have been won — and designs the next race to find out.
The enemy this book is written against
Name it cleanly so the rest of the chapters do not have to keep re-litigating the mood:
- Ship-the-winner analytics — the tournament ends at the trophy.
- Surface cloning — “more like B” as a substitute for isolation.
- Dashboard amnesia — learnings that live only in a vendor UI die with the campaign and the person.
- Heat-as-truth — treating engagement as authority over what deserves to remain in the canon.
None of those enemies is “using data.” The enemy is using data as a selection ritual while starving the organisation of an interpretable model of idea-space.
What changes when experiments teach
A successful test should leave more behind than a flag on a variant ID. It should leave an observation with dates and exposure; candidate mechanisms, not one forced story; concepts implicated at a grain you can join later; confounders you refuse to pretend you controlled; a falsifying next test; and a promotion decision (or a deliberate non-promotion).
That residue is the beginning of a semantic experiment graph: a learning system where stimuli are decomposed into named semantic and surface variables, outcomes file as receipts onto those variables, and graph structure proposes the next informative probe.
Editorial kill list
- Not invented lift percentages or conversion folklore.
- Not one-post causality cosplay.
- Not the quote satellite architecture (sibling work).
- Not the full inbound radar as personal intelligence (later sibling).
If you already run A/B programmes and still cannot answer “why did it win, and what should we isolate next?”, you are the reader. If you only want a bigger uplift percentage with no interest in mechanisms, the rest will feel like friction. That friction is the product.
Key Takeaways
- Conventional A/B is excellent at crowning objects and weak at naming mechanisms.
- Artefact-level scores confound the variables you actually need to learn from.
- Generative cheapness makes mechanism design, not variant count, the scarce skill.
- The durable asset is not the winning creative; it is the learning system that chose and explained it.
Claim-Grain Stimuli Make Response Addressable
Chapter 1 blamed the wrong unit of analysis. This chapter replaces it: when the public unit is roughly one meaning-complete idea, heat finally has somewhere true to land.
If the stimulus contains fifty ideas, the market’s response is an average over fifty ideas plus wrapper plus mood. You cannot ask the score which thought did the work any more than you can ask a blended smoothie which fruit was ripe. Marketers still try. They highlight a sentence they liked. They guess. They run another multi-idea package and guess again.
Claim-grain stimuli change the measurement geometry. When the public unit is roughly one meaning-complete idea — one tension, one mechanism, one ask — heat attaches to roughly one concept. That is not mystical. It is instrument design. You arranged for the artefact to coincide with the idea, so platform telemetry finally has somewhere true to land.
Relational resolution, not “more clips”
An earlier article in this series named the deeper move behind “cutting up a pillar.” The traditional model is hierarchical: expensive parent, lesser descendants. The useful model is a graph: a rich source yields meaning-complete units; each unit can be resolved against a wider corpus; new joins become first-class. Descendants are not automatically smaller. They can be new acts of cognition over the original.
The mechanism is relational resolution: the increase in precision with which a meaning-complete fragment can be positioned, interpreted, and joined. A fifty-idea ebook can only be related broadly — “it develops three frameworks.” Fifty coherent claims can each contradict, extend, evidence, or collide with something specific. Decomposition does not merely reduce the object. It increases semantic addressability.
Semantic refraction is the memorable metaphor for that separation: a rich source passes through different lenses and yields distinct valid directions rather than a shrunken average. That doctrine is already published; this book borrows the grain, not the whole argument. What matters for experiments is the measurement consequence: addressable meaning makes addressable response.
Decompilation produces testable units
You do not get claim-grain by asking a model to “summarise harder.” Summary compresses many ideas into fewer ideas — usually lossy. Extraction isolates one meaning-complete idea faithfully. Interpretation makes significance explicit. Relational enrichment joins it. Media compilation renders it for a moment and channel. Experiment design needs the last four, not the first.
Semantic decompilation is the method that recovers those pieces as design elements — claims, roles, dependencies — rather than as chatty paraphrases of a blob. Again: the method is owned elsewhere. Here we only need the industrial output: tractable objects of cognition small enough to inspect, rich enough to connect, still attached to provenance. Those objects are what you can preregister as experimental variables.
Platforms measure artefacts. You can arrange for artefacts to coincide with concepts.
Why “concept-level telemetry” was never on the vendor menu
Analytics platforms measure artefacts because artefacts are what you hand them. They cannot sell you concept-level heat if your concepts never appear as stimuli. When you ship claim-sized units that carry edges into your own graph, something new becomes possible:
- heat on a quote or card attributes toward a claim and upward to a concept;
- several renderings of the same claim let you compare presentation effect (variance) against idea effect (shared response shape);
- null response becomes informative at the same grain as success.
That is not a guarantee of causal purity. Timing still exists. Algorithms still exist. Audiences still change. It is a massive reduction in confounding relative to multi-idea pillars. Good experimental design is mostly honest confounding reduction, not magic.
Substrates this book assumes
Two sibling architectures
Attention-native publishing treats the interrupt as a scarce petition and long-form as the proof surface. The point for us: you already have a reason to ship small, honest, high-tension units rather than only calendars of bulk.
Quotes without canonical authority keeps exact exhibits addressable with access without citizenship — telemetry and reuse without letting every resonant sentence rewrite the canon. The point for us: you can file heat on working drawings without turning the audience into your truth committee.
If you lack both, you can still run the experiment graph on email subject lines, ad primary texts, or single-claim landing sections. The grain matters more than the medium. The medium only needs to be small enough that “what was tested?” has one honest answer.
The objection
“We already break content into posts. Why is this different from a content calendar?”
A content calendar multiplies surfaces for coverage. Claim-grain instrumentation multiplies addressable bets. Coverage asks: did we show up everywhere? Instrumentation asks: can we say which idea the world pushed back on? If your fragments are still multi-idea packages with no preregistered variables, you are calendaring, not experimenting.
“Atomic posts feel thin.”
Thin is a proof-surface problem, not an experiment problem. The long argument still has to exist somewhere falsifiable. The stimulus is allowed to be incomplete for attention; it is not allowed to be multi-headed for measurement. Completeness belongs to the set, not every container.
Key Takeaways
- Multi-idea artefacts make heat unattributable by construction.
- Meaning-complete claim grain is a measurement architecture, not only a publishing style.
- Borrow refraction and decompilation for units; use them here to attach response.
- Arrange artefacts to coincide with concepts, then platform scores become concept telemetry candidates.
Two Compositions and the Outcome Receipt
Claim-grain is necessary but not sufficient. If you never name what you tested, every stakeholder invents a different hero after the fact.
Once the stimulus is small enough to address, the next failure is sloppy description. Teams ship a claim-sized post and still cannot learn because they never named what they were testing. After the fact, every stakeholder invents a different hero: the hook, the founder story, the colour, the topic of the week. The result is three incompatible “learnings” from one number.
The fix is bureaucratic in the best sense. Preregister two compositions for every stimulus, then attach a dated receipt that is allowed to remain uncertain.
Semantic composition
This is what the idea is made of — the interpretable sparse vector:
| Slot | What you name |
|---|---|
| Concepts / frameworks | Which named ideas are live? |
| Claim and mechanism | What must be true for this to work? |
| Tension or contradiction | What belief does it collide with? |
| Human implication | Why should a person care in their body? |
| Evidence type | Historical analogy, field note, formal result, parable? |
| Adjacent IP | What neighbouring canon does it touch? |
If you cannot fill this table without waving your hands, you are not ready to learn from the result. You may still publish — attention systems do that — but you should not pretend the metric will teach.
Surface composition
This is the presentation and delivery envelope:
| Slot | What you name |
|---|---|
| Hook | First collision line or visual entry |
| Length / density | How much work the reader must do |
| Image / layout style | Rendering choices |
| Structure | Diff, story, list, proof chain |
| CTA | What action, if any |
| Channel and timing | Where and when the petition lands |
Surface variables are not “less pure.” They are different. Conflating them with semantic variables is how you conclude that monochrome images cause engagement when the idea may have been doing the carrying.
The outcome receipt
A receipt is not a vanity row. It is an immutable observation plus a revisable interpretation layer. The honesty posture matches ranked candidates with receipts elsewhere in soft-data work: you keep the system from laundering uncertainty into fake certainty.
Ledger fields
- stimulus ID + composition snapshot
- published_at, channel, audience slice
- exposure
- outcomes by type
- relative band (labelled practitioner, not gospel)
Interpretation fields
- confounders
- candidate explanations ranked
- concepts to credit or debit as priors
- next falsifying test
- promotion decision
One run produces candidates. Several related runs produce hypotheses. Controlled isolation produces patterns. Cross-channel recurrence earns something like marketing canon. Skipping steps is how folklore gets a logo.
Outcome types are not interchangeable
A like is not a consulting enquiry. A save is not a click. A high-quality disagreement in comments may be more valuable than polite agreement. If your heat function blends them into one score, you will optimise for the cheapest dopamine the platform sells.
Heat should be indexable at least by heat[concept, audience, channel, objective, format, time]. A concept can be hot for thought leadership on a professional network and cold for cold email. That is not inconsistency. That is reality having dimensions. Single-score heat will eventually mislead you; treat that as a design bug, not a philosophical mystery.
Silence is a result
If you only file heat when numbers spike, you are sampling on the dependent variable. A stimulus that should have landed and did not is evidence. File the null with the same schema. Later chapters will need those negatives when convergence tries to send you back into a crowded, fatigued cluster.
Worked miniature (illustrative)
Two posts share a concept — “judgment becomes scarce as implication-walking gets cheap” — but differ in surface: one uses a historical doorway and a reflective close; the other uses a punchy list and a demo CTA. If the first draws deep comments and the second draws shallow reactions, a blended score might call them “similar.” Separate receipts will not. They will suggest different objectives and different next tests. That suggestion is still a candidate, not a law. Candidates are enough to design the next race.
Promotion without theatre
Receipts are not a museum of screenshots. They are the input to a promotion ladder you will meet again in Part III: one careful result becomes a candidate explanation; several related results become a hypothesis; controlled isolation becomes a pattern; cross-channel recurrence earns marketing canon. If your culture skips from screenshot to “we now know,” the receipt schema was theatre. The schema only works if someone is willing to write the sentence “we will not conclude X from this run.”
Objection
“Preregistration slows marketing down.” It slows mutation theatre. It speeds learning. Writing two short tables takes minutes; re-learning the same superstition every quarter costs headcount. If a team refuses to name variables before shipping, they have already chosen superstition with traffic.
“We don’t have full funnel data.” Use the outcomes you have. Partial receipts beat theatrical certainty. The crime is not incomplete telemetry. The crime is incomplete description of the stimulus.
Key Takeaways
- Preregister semantic composition and surface composition separately.
- File dated receipts with multi-type outcomes, confounders, and ranked candidates.
- Do not blend unlike outcomes into one heat idol.
- Silence and failure are first-class observations.
Heat Is a Nudge, Not the Truth
Once heat can project onto concepts, teams recreate the old failure in a smarter costume: they let engagement become truth authority. Do not.
By now the machinery is almost dangerous. You have claim-grain stimuli, named compositions, and receipts. Heat can project onto concepts. That is exactly the moment teams let the crowd become the janitor.
What heat is allowed to be
Heat is a multi-dimensional prior about market readiness and response surface. It answers questions like which concepts currently draw lean-in from which audiences; which formats carry a concept without killing it; which objectives a concept serves; which ideas look fatigued even if they still “feel important” internally.
Heat does not answer whether a claim is true; whether a framework deserves canon membership; whether a contradiction should be collapsed; whether the janitor should delete a page because it is unpopular.
That split is Nudge Doctrine applied to outbound telemetry. A nudge biases a judgment layer. It does not decide. In retrieval systems, the doctrine demotes fuzzy signals — similarity scores, mention counts, graph rank — from oracle to prior so deterministic structure and one clear judgment step keep authority. Marketing heat is the same class of signal: frequent, soft, useful, and corrupt if promoted.
Engagement nudges. It does not command the canon.
The audience is a bad janitor
Audiences optimise for lean-in under platform incentives. They are not convened as an epistemic committee. If heat ranks authority, you build enthusiasm inflation: true-but-quiet structure loses to spicy-but-thin hooks; the corpus drifts toward what travels, not what holds. Over a year that drift is indistinguishable from a brand losing its spine while “following the data.”
File heat on quote pages and concept nodes as dated observations. Aggregate upward as advisory metrics. Treat them the way a good navigator treats embedding recall: a hint you can ignore when the walk says otherwise.
Multi-dimensional or bust
Chapter 3 introduced indexing heat by concept, audience, channel, objective, format, and time. Enforce it operationally or the prior becomes a lie.
- A concept is hot for senior operators on a professional network and cold for consumer social. Do not average that into “medium heat.”
- A post produces disagreement that improves the argument. That may be high-value heat under a “stress-test the canon” objective and low-value under a “book demos” objective.
- Last quarter’s spike is not this quarter’s prior without decay. Chronological stacking gives you decay for free if you stop overwriting history with a single lifetime score.
Exploration budget: keep the sensor a sensor
If the experiment graph only proposes high-predicted-heat hybrids, you will overfit to crowd-pleasers; never detect readiness for alien-but-important ideas; and launder prediction into destiny.
Borrow the “alien signal” instinct from inbound sensing and mirror it outbound. Reserve a small, explicit budget of low-predicted-heat stimuli — not as wasted content, as calibration and exploration. Some will fail. That is the point. Failures update the map. A map updated only by successes is a sales brochure.
Anti-causality posture
One outlier does not prove a mechanism. A cluster of related receipts raises a hypothesis. Controlled descendants isolate variables. Cross-channel recurrence earns pattern status. Jumping steps is how “our data shows” becomes a theological phrase.
Wall poster
Ranked candidates with receipts, not asserted causality.
When someone asks “did the poor image cause the win?” the honest answer is almost always: it is a ranked candidate among several, and the next tests should try to kill the wrong ones.
Practical filing rule
When a result lands, write three sentences before you celebrate:
- What heated — concept names, not post IDs.
- What confounded — the list you hate writing down.
- What we will not conclude — the overclaim you are refusing.
That third sentence is the difference between a learning culture and a storytelling culture. Storytelling cultures always “knew it would work.” Learning cultures keep a written record of the explanations they declined.
Objection
“If we don’t let heat decide, what is the point of measuring?” Measuring updates priors that change which experiments you run and how you allocate scarce attention. That is enormous. It is not the same as letting the crowd edit doctrine. Navigation is not legislation.
“Our executives only read one score.” Then your executive interface is wrong. Give them a small panel: top heating concepts, top cooling concepts, one exploration note, one promotion candidate. If the organisation can only metabolise a single vanity number, no graph will save it — but you still should not corrupt the underlying receipts to match the pathology.
Key Takeaways
- Heat is a multi-dimensional advisory prior about response, not truth.
- Never let engagement rank canon authority or drive janitorial deletes.
- Index heat by context; decay it over time; separate outcome types.
- Keep an exploration budget so the system remains a sensor.
Decompose the Outlier Before You Clone It
Part I built the instruments. Part II puts a real outlier on the bench — not to reenact it, but to name its variables.
The case is a professional-network post that used Einstein’s thought experiments as a doorway into a claim about modern reasoning models. Relative engagement landed far above the author’s usual baseline — reported around twenty-four times baseline as a practitioner observation, not as a platform-audited study. The image was poor monochrome, almost a mis-post. That combination is why the case is useful. A pretty image that wins confounds everything. A weak image that still travels makes “the wrapper did it” a weaker candidate and forces attention onto the semantic vector.
This chapter does not claim the post proves causality. It preregisters what the post was made of so later tests can isolate pieces. If you skip decomposition, your next move will be superstition: more Einsteins, more monochrome, more of whatever the dashboard can see.
What the post actually argued
The arc was one mechanism increasing in consequence, not five AI slogans stapled together:
- Einstein often lacked apparatus for immediate verification; he ran implications forward in thought.
- The lab verifies; discovery bottleneck can live in implication-walking.
- Reasoning models cheapen exhaustive implication tracing.
- Scarcity migrates to premise selection and answer recognition — judgment.
- The cliché “AI makes us all Einstein” is rejected.
- The closing question is identity-level: what are you getting better at — running experiments, or knowing which ones to run?
Public history supports the doorway without turning this into a physics textbook. Thought experiments are a recognised mode of reasoning with imagined scenarios.4 The 1919 eclipse observations became a famous early public test of light-bending predictions long after the theoretical work.5 The marketing structure is: historically legible mechanism → present discontinuity → scarcity migration → anti-hype guardrail → personal test. Cognitive time travel is the nearby framework for implication-walking as compressed access to future work states.
Preregistered semantic composition
| Slot | Named contents (illustrative register) |
|---|---|
| Concepts | Gedankenexperiment lineage; implication-walking; bottleneck migration; judgment scarcity |
| Claim / mechanism | Formalisation of implications is no longer the scarce step; choosing and recognising is |
| Tension | Collides with “more model power makes everyone a genius” |
| Human implication | Your skill stack may be training the wrong muscle |
| Evidence type | Historical analogy carrying a precise mechanism |
| Adjacent IP | Intention / goal-formation cluster; anti-hype patterns |
Preregistered surface composition
| Slot | Named contents |
|---|---|
| Hook | Patent office / no lab doorway |
| Length | Long-form feed essay |
| Image | Poor black-and-white (candidate negative or neutral presentation effect) |
| Structure | Progressive argument → scarcity shift → cliché rejection → identity question |
| CTA | Soft link; reflective close is the real ask |
| Channel / timing | Professional network; calendar confounders |
Candidate explanations (ranked, not crowned)
- Concept density and internal alignment — every paragraph advances one scarcity-migration idea.
- Historical genius as vehicle — recognition tax is low; doorway does work.
- Constraint-shift payoff — readers leave with a portable reframe.
- Anti-hype rejection — credibility via boundary.
- Identity-level close — tech observation becomes self-test.
- Timing / topic salience — models-in-the-news confounder.
- Image magic — weak candidate given quality, still not impossible.
Confounders to keep on the receipt
Distribution quirks, audience composition that week, algorithm mood, prior-post priming, and simple luck of a topic wave. If your culture cannot list confounders without feeling unloyal to the win, it is not ready for an experiment graph.
Wrong lessons vs right posture
The dashboard wants: “black-and-white images work,” “always use famous scientists,” “long posts always beat short,” “reflective questions always convert.” Those are surface clones waiting to happen.
The almost-mis-post is better data than the polished twin that confounds the lesson.
Treat the outlier as a high-value observation with a dense semantic vector. The poor image is a gift: it lowers the prior on pure presentation heroics. Your job is not to reenact the post. Your job is to design descendants that try to kill incorrect explanations. Chapter 6 is that programme: hold concept and swap carrier; hold carrier and swap concept; hold structure and vary both. Three cheap posts later you have isolated a variable no ad platform can name.
Notice what decomposition refuses to do. It refuses to declare a single winner mechanism from one run. It refuses to let the relative engagement band — even a dramatic practitioner band like “far above baseline” — become a causal certificate. It refuses to treat Einstein as interchangeable with “any famous person.” Those refusals are the product. Without them you will industrialise the wrong lesson at scale.
Objection: “If we don’t know why it won, how can decomposition help?” Decomposition does not invent the cause. It names the variables so the next three posts can isolate them. Without names, you cannot design isolation. Without isolation, you will reenact superstition with higher production values.
Key Takeaways
- Use outliers as benches for preregistered decomposition, not as templates for cloning.
- Separate semantic vector from surface envelope before you celebrate.
- Rank candidate explanations; keep confounders visible.
- Weak presentation that still travels is unusually informative — still not proof.
Controlled Descendants: Carrier, Concept, Structure
Mutation theatre reenacts winners. Controlled descendants try to kill wrong explanations. The examples below are illustrative, not measured results.
Chapter 5 left a dense outlier on the table. The conventional next step is mutation theatre: more historical-genius posts, more monochrome, more patent-office vibes. That is how you convert a lucky measurement into a brand costume.
The semantic experiment graph asks for controlled descendants — new stimuli designed to isolate a variable class. These examples are illustrative. They are not reported results. Do not read them as measured wins.
Design A — Hold concept, change carrier
Hold: scarcity migration of judgment under cheap implication-walking.
Change: the historical doorway.
Illustrative carriers: Darwin waiting on evidence structures that had not yet arrived; Turing specifying machines ahead of hardware; Ramanujan asserting results whose formal apparatus lagged; a planner reasoning through futures instruments could not yet verify.
If heat holds (shape, not number): the concept is doing real work beyond Einstein celebrity.
If heat collapses: the famous person may have been the asset.
Design B — Hold carrier, change concept
Hold: Einstein / patent-office / thought-experiment doorway.
Change: the payload concept.
Illustrative payloads: formalisation-cost collapse; “the question is source code” / prompt-as-source framing; goal formation as the scarce reagent; wiki-as-kernel compounding once formalisation is cheap.
If heat holds: the doorway is a general-purpose vehicle.
If only the original payload worked: you may have been measuring Einstein-plus-that-specific-reframe.
Design C — Hold structure, vary both
Hold: the rhetorical machine — set up a popular belief, contradict it with a mechanism, close on an identity-level question.
Vary: carrier and concept together.
Illustrative structure holds
- Popular belief: “AI removes the need for specialists.” Diff: specialists become premise-choosers. Close: which judgments are you practising?
- Popular belief: “Faster drafts mean better marketing.” Diff: faster drafts increase the cost of bad selection. Close: what did you refuse to ship?
What you learn: whether the structure is a reusable mechanism independent of Einstein and independent of one concept family.
Why three posts beat a hundred clones
Clones re-estimate the same confounded package. Isolation designs give you a local experimental programme. After a small set of descendants you can often demote two of the seven candidate explanations from Chapter 5 and promote one or two for further tests. That is learning. Reproducing the trophy creative is collecting souvenirs.
This is the marketing twin of two prior doctrines. The Author’s Attention argued for evaluating whether the system’s choosing capacity improved, not only whether a ranking moved. Replay-driven design evolution treats early artefacts as laboratory glassware: the durable assets are the harness and the history.
Descendant brief template
Parent stimulus ID: ... Held constant: [concept | carrier | structure | rendering] Deliberately varied: ... Predicted loser explanations this should hurt: ... Receipt fields to watch: ... Promotion path if pattern repeats: ... Exploration? (yes/no)
If the brief cannot name what it is trying to hurt, it is not an experiment. It is content.
Sequencing and rendering isolation
Practical order: carrier swap first when celebrity effects are loudest; concept swap second once you know whether the doorway generalises; structure hold third when you suspect the rhetorical machine is the real product. Between each, file receipts and update the ranked candidate list.
When capacity allows, ship several renderings of the same claim: different image, different length, different hook, same semantic composition snapshot. Variance across renderings is a presentation-effect band; shared response shape is an idea-effect candidate. Still ranked, still confounded by time — but cleaner than anything artefact-only analytics can say.
What changes in the Monday meeting
After you adopt controlled descendants, the Monday meeting stops asking only “what won?” and starts asking “which explanation did we hurt?” That question is uncomfortable for teams raised on trophy screenshots. It is also the only question that compounds. A wall of winners without demoted explanations is a highlight reel, not a learning system.
Keep the emotional texture honest: isolation can feel slower than cloning. It is slower at producing lookalikes. It is faster at producing knowledge. If leadership measures the marketing function solely by volume of variants shipped, you will need an explicit isolation budget written into the operating rhythm (Chapter 9) or the graph will starve.
Objection
“Isolation is academic; we need pipeline volume.” Volume without isolation is a content factory with a superstition layer. Define a fraction of tests as isolation descendants or exploration probes. Without that fraction, the graph never updates.
“What if all three descendants fail?” Then the outlier may have been timing, network effects, or an irreproducible bundle — cheaper than building a brand costume on a one-off.
Key Takeaways
- Design descendants that hold concept, carrier, or structure — one class at a time.
- Label illustrative programmes as illustrative until measured.
- The goal is to kill wrong explanations, not to reenact the winner.
- Harness and history compound; trophy creatives perish.
Local Gradients and Graph-Guided Next Tests
After isolation designs return receipts, stop asking what the last winner looked like. Ask where independent heat points.
A dashboard winner is a label on an object. The experiment graph wants a sparse, named vector over concepts and mechanisms:
gedankenexperiment: high bottleneck-migration: high human-judgment: high anti-hype-reversal: medium famous-person-entry: unknown (being isolated)
Dense embeddings can tell you what resembles the winner. They will not tell you what the dimensions mean. Typed concepts with edges give you a basis that is legible by construction. Heat projects onto those nodes. Edges carry some of that heat as a prior — never as legislation.
Walk the heat
Take independently hot nodes. Follow typed edges. Rank what is most pointed-at. This is mechanically close to convergence operations already used when multiple probes meet: independent routes bloom an intersection worth inspecting. Marketing inherits the same geometry with heat as a weight.
Under-rendered concept
Heat points at a node that exists but has few public stimuli. Next test: ship a clean claim-grain stimulus for the audience that generated the heat.
Empty space
Edges converge where no page yet exists. Candidate missing pillar — not automatic truth. Probe, then promote only after the receipt ladder.
There is a third landing: accident cluster — a temporary news wave gluing unrelated heat. Confounders on receipts are how you notice.
Cross-post convergence
One post is a weak gradient. Several posts that implicate overlapping concepts start to draw a field. That is why filing receipts against concepts beats filing them against campaign IDs. Campaign IDs die. Concepts persist. Over months you get something no single A/B tool shows: which regions of your idea-space the world can currently hear.
What a successful experiment must leave behind
Make this checklist non-negotiable. If a test does not produce these artefacts, it was entertainment with metrics.
- Observation — dated, with exposure and outcome types.
- Candidate mechanisms — ranked.
- Concepts implicated — named nodes, not vibes.
- Confounders — written while memory is fresh.
- Falsifying next test — the descendant that would hurt the leading explanation.
- Adjacent concepts — where edges point.
- Suggested hybrids — only if convergence supports them.
- Promotion decision — raw / hypothesis / pattern / canon — or explicit reject.
The rig still matters. The durable asset is the accumulating theory the rig improves.
Worked sketch (illustrative)
Suppose heat is independently high on human judgment, problem framing, intention compiling, and goal formation. Edges point toward a shared neighbour: “execution is abundant; choosing what deserves execution is scarce.” That neighbour might be an existing page under-rendered on the professional network; a latent concept needing a name; a strong ebook thesis; or a temporary hype rhyme.
The graph does not decide. It tells you where to inspect. Humans and explicit promotion rules decide what becomes doctrine. Signal Case Queue thinking helps: bound the ambiguous convergence into a case with evidence, options, and a decision, rather than letting a pretty graph picture spend headcount automatically.
Gradient, not prophecy
A local gradient tells you the direction of informative travel, not the guaranteed prize. If three hot nodes point at empty space, the correct next move is a carefully framed probe — not a six-week ebook commissioned on vibes. The probe’s job is to test whether the empty space is a real demand surface or a mirage built from correlated confounders.
From experimental search to next week’s queue
Translate the gradient into a queue the team can run without philosophy seminars. After each harvest cycle, produce three short lists: (1) under-rendered concepts with heat and few recent stimuli; (2) empty-space candidates with the edges that pointed there; (3) accident clusters to ignore this week. Assign each item a proposed test type — isolation, hybrid, exploration, or ordinary shipping. That is experimental search as operations, not as a slide.
Embeddings still have a job: they can cluster surfaces and find lookalikes when you need production help. Just do not let resemblance replace the typed walk. Resemblance finds cousins of the winner. The graph finds mechanisms and holes. You need both; only one teaches the compiler why the race may have been won.
Objection
“This sounds like we need a full wiki before we can market.” You need named variables and a place to file receipts. That can start as a disciplined spreadsheet and a concept list. The wiki is how it compounds at organisational scale. Waiting for perfect graph infrastructure is another way to stay in mutation theatre.
“Convergence will send us into jargon.” Then your surfaces are failing, not your concepts. Heat on a jargon bomb may be a mandate to re-render, not to ship more jargon.
Key Takeaways
- Project heat onto a sparse named concept basis; walk edges for next tests.
- Distinguish under-rendered nodes, empty space, and accident clusters.
- Require a leftover checklist so tests update theory, not only trophies.
- Convergence nominates; promotion rules and humans legislate.
Cross-Heat Hybrids and the Marketing Canon
Isolation cuts variables apart. Cross-heat hybrids put the right ones back together — as walks from receipts, not brainstorms with better fonts.
If Chapter 7 is navigation, this chapter is constructive geometry: how new public theses get proposed by the field you measured, and how some of those theses earn a place in a marketing canon that outlives any campaign.
Hybrids are not brainstorms
A brainstorm starts from cleverness. A cross-heat hybrid starts from receipts:
Hot node A Hot node B Shared or bridging neighbour N → draft stimulus that makes N explicit
Examples in the neighbourhood of the Einstein outlier (illustrative combinations, not measured winners):
- Gedankenexperiment × Intention Compiler — the premise is source code for the future the model will construct; bad premises compile fluently into confident waste.
- Cognitive time travel × goal formation — many futures are now cheap to access in simulation; value lives in which future deserves access.
- Formalisation-cost collapse × wiki-as-kernel — once models can formalise quickly, the compounding asset is the kernel those formalisations write back into.
- Thought experiment × synthetic futures — reasoning models are laboratories for consequences, not oracles of random events.
Each hybrid is a candidate stimulus with its own preregistered compositions. It still gets a receipt. It still faces isolation descendants if it spikes. Hybridisation is how the graph generates informative novelty; it is not a free pass past the method.
Why heat-guided hybrids beat “more content like the winner”
Surface similarity search finds neighbours in embedding space: posts that look like winners. Typed heat finds neighbours in mechanism space: ideas that share causal load with what worked. Those sets overlap sometimes and diverge often. The diverge cases are where the experiment graph earns its keep — proposing a next pillar the resemblance engine would never suggest because the wording has not existed yet.
The promotion ladder
Campaign dashboards are graveyards of untransferred learning. When the person leaves, the “knowledge” leaves. Filing against concepts moves learning into a space that can compound.
| Stage | What earns it | What you store |
|---|---|---|
| Raw observation | One receipt | Ledger entry |
| Candidate explanation | One careful read | Ranked hypotheses |
| Marketing hypothesis | Several related receipts | Testable statement |
| Application pattern | Controlled / repeated tests | Playbook fragment |
| Marketing canon | Cross-channel recurrence + human sign-off | Generalised doctrine |
This is the Proposal Compiler’s feedback posture at finer grain: responses update frameworks and positioning, not only win/loss tallies. Newsjacking with a canon already treated engagement as signal filed back; concept-addressed receipts give that signal a stable home.
What must never auto-promote
- A single viral spike into canon language.
- A high-like / low-quality-attention pattern into “what we believe.”
- A celebrity carrier effect into a claim about universal concept strength.
- Anything that would let heat reorder truth lanes in the IP wiki.
Canon is curated judgment informed by telemetry. Telemetry alone is a mood ring.
Hybrid brief
Parents: [hot concepts A, B, …] Bridge: [neighbour N or empty-space name] Why heat suggests this (receipt IDs): ... Semantic composition: ... Surface composition: ... What would falsify the hybrid: ...
If you cannot point at receipt IDs, you are brainstorming. Brainstorms are allowed — just do not call them graph-guided.
Worked institutional scene
A team runs twelve isolation and hybrid probes over a quarter. Three patterns survive: historical mechanisms outperform historical inspiration porn for their operator audience; anti-hype guardrails improve comment quality on ambitious AI claims; identity-level closes outperform generic CTAs for conversation objectives — not necessarily for click objectives.
Those three sentences are marketing canon candidates. They are more valuable than any single creative file. They survive tool changes. They onboard new marketers. They give the experiment graph something to protect when a new platform metric tries to reassert single-score tyranny.
Institutional memory wearing a marketing hat
Most organisations already believe institutional knowledge matters for product, ops, and sales. They still treat marketing learning as disposable campaign exhaust. That asymmetry is expensive. When heat receipts live in concept-space, a new marketer inherits mechanisms instead of a folder of “best performing creatives” whose context has evaporated. When heat stays trapped in a vendor dashboard, the company re-buys the same lesson every year with new software logos.
This is also why promotion must stay human-gated. Auto-promoting viral language into canon is how brands become platforms’ echo chambers. The graph proposes; the canon editor refuses enthusiasm inflation. Refusal is a first-class output of the system, not a failure of analytics.
Objection
“Canon will ossify our creative.” Bad canon will. Good canon states mechanisms and still demands isolation when context shifts. “Anti-hype helps our audience” is a prior, not a forever law. Re-test when audience or channel changes. The point of canon is to stop relearning the same lesson every Monday — not to stop learning.
Key Takeaways
- Build hybrids by walking hot nodes, not by mutational cleverness alone.
- Promote learnings through a staged ladder into a marketing canon.
- Never let a single spike rewrite doctrine.
- Concept-space memory outlives campaigns, platforms, and people.
Operate the Semantic Experiment Graph
The last chapter is not a new framework. It is the same doctrine under operational load: ledger, graph, canon, budgets, and a Monday start smaller than you think.
Three layers, three jobs
Performance ledger
Immutable-ish observations. What shipped, exposure, outcome types. Keep it boring and complete.
Experiment graph
Derived interpretation: concepts, mechanisms, candidates, confounders, edges, heat priors. Soft and versioned.
Marketing canon
Patterns that survived the promotion ladder and a human sign-off. Short enough to teach; strict enough to resist mood.
If you collapse all three into a dashboard screenshot, you are back in Chapter 1.
A weekly loop (shape, not dogma)
- Harvest — file receipts for everything that shipped, including silence.
- Update priors — adjust concept heat multi-dimensionally; decay old spikes.
- Inspect convergence — under-rendered nodes, empty space, accident clusters.
- Queue tests — isolation, hybrids, ordinary shipping, exploration probes.
- Preregister — semantic + surface tables before publish.
- Ship — claim-grain where measurement matters.
- Promote or refuse — explicit decisions, not vibes.
Attention-native discipline still applies at the petition layer: not everything becomes an interrupt. The experiment graph does not argue for more noise. It argues that the noise you do make should teach.
Budgets that keep the system honest
- Isolation budget — a defined fraction of tests must hold a variable class constant.
- Exploration budget — a defined fraction may be low-predicted-heat.
- Canon budget — very few items promote; promotion is expensive on purpose.
- Interrupt budget — heat does not override scarcity of audience trust.
Without these budgets, optimisation collapses to “ship what looks like last week’s winner.”
Roles (even in a team of one)
Stimulus author writes claim-grain units with preregistered compositions. Receipt clerk files outcomes and confounders. Navigator walks heat and proposes next tests — must not auto-legislate. Canon editor is the human principal for promotion and refuses enthusiasm inflation.
In larger teams these split. In a team of one they become modes you switch between deliberately. The failure mode is always the same: one mode eats the others.
Boundaries
What this book still does not own
- Quote satellite storage — Quotes Without Canonical Authority.
- Publishing as an active sensor in a full inbound/outbound intelligence loop — later work under that name owns the causal architecture.
- Future siblings — Executable Worldview, Intent-Conditioned Task World, Orientation Capital, Institutional Memory Not Cognition — may extend how organisations hold what experiments teach. Not prerequisites for starting Monday.
Relational arity and prompt-as-source remain ambient constraints: tags are not joins; disposable surfaces are not the durable source of learning. Cite them when you instrument; do not rebuild them here.
Start smaller than you think
You do not need a perfect graph database to begin. You need:
- one outlier decomposed;
- three controlled descendants queued;
- a receipt schema used twice;
- one promotion refusal (to prove you can refuse);
- one exploration probe you did not expect to win.
Do that for a month and you will feel the difference between selection rituals and experimental search. The graph can get richer later. The honesty has to show up first.
A/B testing crowns a winner. A semantic experiment graph teaches the compiler why the race may have been won — and designs the next race to find out.
Continuity check against Chapter 1
Return to the Monday meeting that crowned B and cloned its surface. With a semantic experiment graph in place, the same meeting looks different. B still might ship. The difference is the residue: a preregistered composition, a multi-type receipt, ranked candidates, a named falsifying next test, and an explicit non-conclusion list. The organisation is no longer only selecting. It is updating a legible local gradient over idea-space.
That is the economic claim restated without re-deriving the whole spine: generative systems made variant production cheap; winner creatives remain perishable; causal learning remains scarce; concept-space filing compounds; screenshot culture resets every quarter.
Closing objection
“Isn’t this a lot of process for posts?” It is a little process for an asset class that currently evaporates. Marketing learning is organisational memory wearing a public face. If your company already believes institutional knowledge matters, it is inconsistent to let outbound learning die in a vendor UI. If it does not believe that, no experiment graph will fix the deeper problem — but you will at least see the amnesia clearly.
Start with one outlier, three descendants, two honest receipts, one refusal, and one exploration probe. The tooling can grow. The honesty cannot wait. If you only remember one operating rule, make it this: never let a winner label substitute for a mechanism you can name, isolate, and file.
Key Takeaways
- Operate ledger, graph, and canon as separate layers with explicit promotion.
- Keep isolation, exploration, canon, and interrupt budgets visible.
- Start with decomposition, descendants, receipts, and one refusal — then deepen tooling.
- The winning creative is perishable; the experiment graph compounds.
References & Sources
The evidence base behind every claim — primary research, industry analysis, and technical specifications
Research Methodology
This ebook draws on primary research from standards bodies, independent research firms, enterprise technology vendors, and consulting firms. Statistics cited throughout have been cross-referenced against primary sources.
Frameworks and interpretive analysis developed by Scott Farrell / LeverageAI are listed separately below — these represent the practitioner lens through which external research is interpreted, and are not cited inline to avoid self-promotional appearance.
Industry Analysis & Vendor Research
GrowthBook — What Is A/B Testing? [1]
Randomised assignment controls external factors so differences can be attributed to the change under test
https://www.growthbook.io/blog/what-is-a-b-testing
VWO — What is A/B Testing? [2]
Split testing framed as comparing variants to pick a better performer
https://vwo.com/ab-testing/
Primary Research & Standards Bodies
Interaction Design Foundation — A/B Testing [3]
A/B methods can show behavioural difference without explaining qualitative why
https://ixdf.org/literature/topics/a-b-testing
Wikipedia — Thought experiment [4]
Recognised mode of reasoning with imagined scenarios
https://en.wikipedia.org/wiki/Thought_experiment
Wikipedia — Eddington experiment [5]
1919 eclipse observations as early public test of light bending
https://en.wikipedia.org/wiki/Eddington_experiment
LeverageAI / Scott Farrell — Practitioner Frameworks
The interpretive frameworks, architectural patterns, and practitioner analysis in this ebook were developed through enterprise AI transformation consulting. The articles below are the underlying thinking behind those frameworks. They are listed here for transparency and further exploration — not cited inline, as this is the author's own analytical voice.
Scott Farrell — Semantic Refraction
Relational grain and refraction vs fragmentation
https://leverageai.com.au/wp-content/media/articles/152-semantic-refraction.html
Scott Farrell — Semantic Decompilation
Claim-level recovery of design from source
https://leverageai.com.au/wp-content/media/articles/153-semantic-decompilation.html
Scott Farrell — Attention-Native Publishing
Interrupt budget and publication gate
https://leverageai.com.au/wp-content/media/articles/151-attention-native-publishing.html
Scott Farrell — Quotes Without Canonical Authority
Quote satellite: access without citizenship
https://leverageai.com.au/wp-content/media/articles/156-quotes-without-canonical-authority.html
Scott Farrell — BI Where Wiki Why
Ranked candidates with receipts, not asserted causality
https://leverageai.com.au/wp-content/media/articles/106-bi-where-wiki-why.html
Scott Farrell — Nudge Doctrine
Heat and fuzzy signals as priors, never verdicts
https://leverageai.com.au/wp-content/media/articles/100-nudge-doctrine.html
Scott Farrell — Cognitive Time Travel
Implication-walking and temporal access framing
https://leverageai.com.au/wp-content/media/articles/40-cognitive-time-travel.html
Scott Farrell — The Prompt Is Source
Upstream intent package as durable source
https://leverageai.com.au/wp-content/media/articles/154-the-prompt-is-source.html
Scott Farrell — The Author's Attention
Measure whether choosing capacity improved, not only rankings
https://leverageai.com.au/wp-content/media/articles/89-the-authors-attention.html
Scott Farrell — Replay-Driven Design Evolution
Harness and history compound; early artefacts are lab glassware
https://leverageai.com.au/wp-content/media/articles/145-replay-driven-design-evolution.html
Scott Farrell — Intent Compiler
Convergence: independent routes bloom the intersection
https://leverageai.com.au/wp-content/media/articles/141-intent-compiler.html
Scott Farrell — Signal Case Queue
Bound ambiguous signals into competing cases
https://leverageai.com.au/wp-content/media/articles/143-signal-case-queue.html
Scott Farrell — Proposal Compiler
Responses update frameworks, not only win/loss tallies
https://leverageai.com.au/wp-content/media/articles/32-proposal-compiler.html
Scott Farrell — Newsjacking with a Canon
Engagement as signal filed back into the canon loop
https://leverageai.com.au/wp-content/media/articles/77-newsjacking-with-a-canon.html
Scott Farrell — RAG Metadata Relational Meaning
Unary metadata vs binary relational meaning
https://leverageai.com.au/wp-content/media/articles/155-rag-metadata-relational-meaning.html
About This Reference List
Compiled July 2026. All URLs verified at time of compilation. Regulatory documents and standards specifications are subject to revision — check primary sources for the most current versions.
Some links to academic papers and vendor research may require free registration. Government and standards body publications are freely accessible.