Semantic Experiment Graph — Make Every Marketing Test Teach the Next One
A/B testing is good at crowning a winner. It is weak at teaching the next experiment. When stimuli are decomposed into named semantic and surface variables, and outcomes file as dated receipts onto a concept graph, marketing stops shipping perishable creatives and starts compounding a legible theory of what audiences respond to — and why.
Most teams that “do experimentation” are really doing selection. They ship two versions of a page, subject line, or post. Traffic splits. A metric moves. Variant B wins. The dashboard congratulates itself. Next week the team clones the surface of B — same colour, same celebrity, same length — and calls that learning.
Industry tooling is built for exactly that story: compare variants, pick the better performer, move on.1 Randomisation still matters; it controls seasonality and traffic mix so the change under test can be blamed for the delta.2 The gap is not the existence of controlled experiments. The gap is that the unit of analysis is almost always the artefact, not the idea.
A pillar post that “did well” averages concept, vehicle, structure, rendering, channel, timing, and audience into one score. You learn “this post did well.” You cannot say which thought did the work. Interaction-design literature is blunt on the same point from another angle: A/B methods can show behavioural difference without explaining the qualitative “why.”3
A/B testing crowns a winner. A semantic experiment graph teaches the compiler why the race may have been won — and designs the next race to find out.
Claim-grain is a measurement architecture
Earlier work in this series argued that meaning becomes addressable when sources are decomposed into meaning-complete units rather than summarised into mush. Semantic refraction is the metaphor for that move: a rich source separates into valid directions under different lenses, rather than merely shrinking.4 Semantic decompilation is the method: recover claims, roles, and dependencies as if you were recovering a design, not compressing a blob.5
Those pieces matter here for a different reason. Atomisation is not only a publishing convenience. It is how response becomes attributable. When the stimulus is roughly one claim, heat attaches to roughly one idea. When several renderings of the same claim exist, variance between them is presentation effect and the shared lift is a candidate for idea effect. No analytics vendor sells that separation out of the box, because platforms measure artefacts and most teams never arrange for artefacts to coincide with concepts.
Attention-native publishing already flipped the carrier hierarchy under scarcity: the interrupt is the petition; long-form is the proof surface.6 The quote satellite stores exact exhibits with access without citizenship so they can carry telemetry without rewriting canon.7 This article assumes those substrates and asks a harder operational question: once you can ship claim-sized stimuli, how do you make every result teach the next test?
Two compositions, one receipt
Each published item should carry two separate descriptions before it ships:
| Semantic composition | Surface composition |
|---|---|
| Concepts and frameworks | Hook |
| Claim and mechanism | Length |
| Tension or contradiction | Image style |
| Human implication | Post structure |
| Evidence type | CTA |
| Adjacent IP | Channel and timing |
Then the result — impressions, reactions, comments, saves, clicks, conversations, rejections, silence — attaches to both. Silence is a result. A null response on a concept the world should care about is evidence, not a non-event.
The receipt is not a vanity metric row. It is a dated observation with structure:
- exposure and outcome types (not one blended score);
- audience, channel, time;
- confounders you know you did not control;
- candidate explanations ranked, not asserted.
That is the same epistemic posture as ranked candidates with receipts elsewhere in the canon: you keep honesty by refusing to overclaim a single run.8 Nudge Doctrine applies directly: heat enters as a prior, never a verdict. It must not rank a claim’s authority and it must not drive janitorial consolidation of the canon, or the audience becomes your janitor — and audiences optimise for lean-in, not for true.9
Flagship observation: the Einstein post
One practitioner outlier makes the architecture concrete. A professional-network post used Einstein’s thought experiments as a doorway into a claim about reasoning models: implication-walking used to be scarce; models cheapen that formalisation; the scarce human skills migrate to choosing the premise and recognising the valuable answer. The closing question was identity-level: what are you actually getting better at — running the experiment, or knowing which one to run? Relative engagement landed far above the author’s usual baseline (reported around twenty-four times baseline — a practitioner observation, not a platform-audited study). The accompanying image was poor monochrome — almost a mis-post.
The wrong lesson from artefact analytics is “black-and-white images work” or “Einstein thumbnails perform.” The better candidate set is more interesting: a dense, internally aligned semantic vector — historical mechanism, bottleneck reversal, scarcity migration, anti-hype qualification, personal implication — carried a weak visual wrapper. That is not settled causality. It is a ranked explanation with a useful confounder: presentation effect was probably not the hero.
Public history supports the post’s structural choice without turning this article into physics. Thought experiments are a recognised method of reasoning with imagined scenarios;10 the 1919 eclipse observations became a famous early public test of general relativity’s light-bending prediction long after the theoretical work.11 The marketing lesson is the structure: use a historically legible doorway to carry a precise mechanism, then refuse the cliché ending.
Preregistered decomposition (illustrative schema)
Before shipping descendants, name what the original was made of:
- Semantic: gedankenexperiment lineage; cognitive time travel / implication-walking; bottleneck migration; judgment scarcity; anti-hype reversal; identity question.
- Surface: long-form feed post; historical-genius doorway; weak monochrome image; progressive argument structure; reflective close; professional network channel.
- Candidate explanations: concept density; famous-person entry; constraint shift; reflective close; timing/topic salience.
- Confounders: distribution, audience composition, algorithm mood, day of week.
Controlled descendants, not surface clones
A conventional marketer’s next step is mutation theatre: more historical-genius posts with monochrome images. A semantic experiment graph generates mechanically different descendants that isolate variables:
- Hold concept, change carrier. Darwin waiting decades without the decisive apparatus; Turing specifying machines before hardware; Ramanujan asserting results ahead of formal proof. If heat holds, the concept is doing work. If it collapses, Einstein was the asset.
- Hold carrier, change concept. Keep the patent-office frame; swap the conclusion to formalisation-cost collapse, goal formation, or prompt-as-source. This tests which idea the doorway successfully carried.
- Hold structure, vary both. Keep the “set up popular belief → contradict → personal question” diff structure while swapping concept and carrier. This tests the rhetorical mechanism separately.
Three cheap posts later you have isolated a variable no ad platform can name. That is the difference between ranking the variants you were handed and generating variants that would be informative. Replay-driven design evolution makes the same industrial point one level up: the durable assets are the harness and history, not the first artefact that happened to score.12 The Author’s Attention made the evaluation version of the same correction: do not only ask whether a ranking changed; ask whether the system’s ability to choose improved.13
From winner label to local gradient
Dense embeddings can tell you what resembles a winning post. They will not tell you what the dimensions mean. A concept graph is a sparse, named, typed basis. Heat projects onto nodes. Typed edges propagate some of that heat. When several independently hot concepts point toward the same neighbour, you get the convergence mechanism already used for intent resolution: different routes meet; inspect the intersection.14
Two interesting landings matter for marketing:
- Under-rendered concept: heat points at a node that exists but has few public stimuli. Publish there next.
- Empty space: edges from several hot nodes converge where no page yet exists. That is a candidate missing pillar — market readiness without a named page. Name it carefully; do not let heat mint truth.
Cross-heat hybrids are the generative step. Not arbitrary mashups — walks from active nodes:
- Gedankenexperiment × Intention Compiler → “the premise is source code for the future the model constructs.”
- Cognitive time travel × goal formation → “many futures are accessible; value is which future deserves access.”
A home for marketing learning
Campaign learnings that live only inside a vendor dashboard are organisational amnesia with a UI. They do not survive the campaign, the platform change, or the person who ran the test. Filing receipts against concepts moves learning into concept-space, where it can compound.
Promotion is staged, like any other epistemic system:
- one result → candidate explanation;
- several related results → marketing hypothesis;
- controlled or repeated tests → application pattern;
- cross-channel recurrence → marketing canon.
That is the Proposal Compiler’s feedback posture applied at finer grain: buyer and audience responses improve frameworks, not only win/loss tallies.15 Newsjacking with a canon already treated engagement as signal filed back;16 this article gives that signal an address at concept grain.
Keep an exploration budget. If you only publish what the model predicts will heat up, you stop being a sensor. A small tranche of low-predicted-heat stimuli keeps measurement honest and prevents canon drift toward pure crowd-pleasers.
What this is not
- Not proof that one high-performing post establishes causality.
- Not permission for engagement to decide truth or canon membership.
- Not the quote storage layer — that is the satellite architecture of the prior article in this series.
- Not the full inbound/outbound radar loop as a personal intelligence system — a later piece owns that causal claim under the name Publishing Is an Active Sensor.
Relational meaning is at least binary; metadata alone cannot carry it.17 Upstream prompts and intent packages remain source before disposable generated surfaces.18 Those distinctions stay load-bearing when you instrument marketing: you are not decorating posts with tags. You are attaching response to the same grain at which ideas join.
Start Monday
- Pick one recent outlier (hot or cold). Preregister its semantic and surface composition as if you were about to ship it again.
- Write three controlled descendants — hold concept, hold carrier, hold structure — and label them illustrative until measured.
- Define a receipt schema that separates outcome types and lists confounders and candidate explanations.
- After results land, update concept heat as a prior. Walk edges. Ask where independent heat points. Design the next race for information, not for another trophy creative.
Generative systems made variant production cheap. Causal learning is still scarce. The winning creative is perishable. The experiment graph compounds.
References
- VWO. "What is A/B Testing?" — industry framing of split testing as comparing variants to pick a better performer. https://vwo.com/ab-testing/
- GrowthBook. "What Is A/B Testing?" — randomised assignment controls external factors so differences can be attributed to the change under test. https://www.growthbook.io/blog/what-is-a-b-testing
- Interaction Design Foundation. "A/B Testing" — A/B methods can show behavioural difference without explaining qualitative why. https://ixdf.org/literature/topics/a-b-testing
- Scott Farrell / LeverageAI. "Semantic Refraction." https://leverageai.com.au/wp-content/media/articles/152-semantic-refraction.html
- Scott Farrell / LeverageAI. "Semantic Decompilation." https://leverageai.com.au/wp-content/media/articles/153-semantic-decompilation.html
- Scott Farrell / LeverageAI. "Attention-Native Publishing." https://leverageai.com.au/wp-content/media/articles/151-attention-native-publishing.html
- Scott Farrell / LeverageAI. "Quotes Without Canonical Authority." https://leverageai.com.au/wp-content/media/articles/156-quotes-without-canonical-authority.html
- Scott Farrell / LeverageAI. "BI Where Wiki Why." https://leverageai.com.au/wp-content/media/articles/106-bi-where-wiki-why.html
- Scott Farrell / LeverageAI. "Nudge Doctrine." https://leverageai.com.au/wp-content/media/articles/100-nudge-doctrine.html
- Wikipedia. "Thought experiment." https://en.wikipedia.org/wiki/Thought_experiment
- Wikipedia. "Eddington experiment." https://en.wikipedia.org/wiki/Eddington_experiment
- Scott Farrell / LeverageAI. "Replay-Driven Design Evolution." https://leverageai.com.au/wp-content/media/articles/145-replay-driven-design-evolution.html
- Scott Farrell / LeverageAI. "The Author's Attention." https://leverageai.com.au/wp-content/media/articles/89-the-authors-attention.html
- Scott Farrell / LeverageAI. "Intent Compiler." https://leverageai.com.au/wp-content/media/articles/141-intent-compiler.html
- Scott Farrell / LeverageAI. "Proposal Compiler." https://leverageai.com.au/wp-content/media/articles/32-proposal-compiler.html
- Scott Farrell / LeverageAI. "Newsjacking with a Canon." https://leverageai.com.au/wp-content/media/articles/77-newsjacking-with-a-canon.html
- Scott Farrell / LeverageAI. "RAG Metadata Relational Meaning." https://leverageai.com.au/wp-content/media/articles/155-rag-metadata-relational-meaning.html
- Scott Farrell / LeverageAI. "The Prompt Is Source." https://leverageai.com.au/wp-content/media/articles/154-the-prompt-is-source.html
