Leverage AI

Prediction Receipts: Forecast Heat Without Letting Heat Write the Truth

Preregister multi-lane expectations and a falsifier before you publish. When the outcome arrives, the error updates named priors — not the canon.

Scott Farrell · LeverageAI · Long-form article

The system owner can already generate more candidate posts than a professional social network will ever reward. Quote variants, card framings, hook rewrites, timing shifts — the generative surface is infinite. What remains scarce is an inspectable reason for selecting one candidate, a written expectation of what should happen if the selection was sound, and a record of what the miss taught when the market answered differently.

Most “will this resonate?” tools collapse that problem into a single score. They learn associations among wording, carriers and engagement numbers. They are useful for ranking a queue. They are almost useless for learning, because when the number is wrong you do not know which prior failed — the idea, the surface, the audience slice, the moment, or the distribution path that carried the artefact. The post-hoc meeting then invents a story that flatters whoever is in the room. Next week’s selection is still a vibe.

That is not an argument against forecasting. It is an argument against forecasts that cannot fail in a nameable way. A useful attention forecast is not a prophecy about impressions. It is a reasoned prior assembled from what the graph already knows — which concepts are warming, which framings have receipts, which audiences have been ready before, which carriers are noisy — written down before the artefact leaves the building so the difference between expected and actual response has somewhere to land.

A forecast that cannot name how it could be wrong is not a prior. It is a hope with a decimal.

This article owns the instrument that turns attention forecasts into learning: the Prediction Receipt, written before publication, joined later to an Outcome Receipt, with the difference written back only into named priors. It sits inside the whole-object argument of The Semantic Market Model, and it extends — without rebuilding — the Semantic Experiment Graph’s two-composition preregistration and outcome-receipt machinery.1,2

The reader question is practical: how do you forecast whether an idea will resonate without building a black-box virality score? After this piece you should be able to fill a Prediction Receipt, publish a bounded probe, join the outcome, and update semantic, surface, temporal or audience priors — while holding a hard boundary that response never becomes authority over truth or canon. Four pieces of the instrument carry the weight: the five-vector decomposition that keeps forecasts non-scalar; the paired receipt records designed field by field; the selection mechanism that retrieves candidates against typed edges (including the Cross-Heat Hybrid gate); and the write-back rules that update readiness without rewriting truth.

Honesty constraint A preregistration article that fakes its own outcomes is self-refuting. The complete receipt schema below is designed in full and is in scope. Qualitative mechanism and doctrine are grounded. Ten historical backtests without future leakage, three live predicted-versus-actual pairs, and a repeated-rendering idea-mean versus surface-variance estimate are proposed, not yet run. No example receipt in this article carries a fabricated observed number. Illustrative records are stamped exactly as sibling experiment writing stamps its worked programmes: illustrative, not measured results. Do not read them as measured wins.

Why a scalar cannot learn

After a claim-sized post ships, every stakeholder invents a different hero: the hook, the founder story, the colour, the topic of the week. Three incompatible “learnings” emerge from one number.2 That is the failure the Semantic Experiment Graph already named. Two-composition preregistration exists so teams write, before the run, what semantic composition and what surface composition they believe they are testing — and refuse to conclude what the run cannot support.

This piece adds the missing before-publication half of that honesty: not only “what composition are we testing?” but “what response pattern do we expect, how uncertain are we, and what observation would falsify the forecast?” Without those fields, silence is ambiguous and success is uninterpretable. With them, null can reprice readiness, fatigue or framing rather than dissolve into creator disappointment.3

A candidate does not have a single heat. Response is a relationship — concept, framing, author, carrier, audience, surrounding discourse and moment — becoming observable at an artefact. The durable learning is not “post 412 received 1,730 impressions.” It is a typed account of which ideas moved, for whom, through what doorway, under what temporal conditions. Collapse that into one predicted heat and you destroy the places a forecast can be wrong in a nameable way.

So the instrument begins with decomposition. Before you forecast anything, split the candidate into five named vectors. Each is a place the forecast can fail independently.

Vector What it holds Example failure mode
Semantic Concepts, claim, mechanism, tension, consequence, adjacent IP Right surface, wrong or premature idea; audience hears a different claim than you filed
Authority Originator, confirmer, implementer, messenger, source standing in-domain Prestige carrier did the work; mechanism never landed
Temporal Concept temperature, rising/stable/fading, freshness, saturation, fatigue, proximity to a live case Cold idea in a hot neighbourhood — or a hot idea one cycle late
Surface Hook, length, structure, image, tone, close, channel timing Idea-effect strong historically; this rendering killed recognition cost
Audience Who should recognise the issue, required prior knowledge, expected response type Broad reach from the wrong slice; or quiet among the people who mattered

These are not decorative labels for an embedding. An embedding might find posts that resemble a familiar historical figure. The graph can identify a reusable mechanism: familiar historical doorway, implication-walking, scarcity migration, anti-hype boundary, identity-level question. That mechanism can be tested again with a different carrier. The celebrity cannot.

Expected outcomes must stay multi-lane for the same reason. A post can have broad reach and weak intellectual response; low reach and exceptional comments from the right people; strong disagreement and high conceptual importance; high clicks from sensational framing; low visible engagement and important private messages; high audience resonance and no evidence of truth. Lanes that matter for learning include at least:

A single scalar may still be derived for scheduler convenience. Learning must not live in that scalar. A like is not a consulting enquiry; outcome types are not interchangeable; single-score heat will eventually mislead you — treat that as a design bug, not a philosophical mystery.2

Two-composition preregistration already forces the semantic table and the surface table to be fillable without hand-waving. If you cannot fill them, you are not ready to learn from the result.2 The Prediction Receipt adds the forecast layer on top of those compositions: expected bands per lane, uncertainty, and a falsifier written in the same ink as the claim. That is the extender’s job — not to re-teach the experiment graph, but to make the before-publication forecast as first-class as the after-the-fact receipt.

One more distinction before the artefact itself. Predicting significance of an external event, predicting attention to your own probe, and generating informative probes are related capabilities that must not be collapsed. Inbound systems can recognise that a world configuration now resembles a consequential motif in the owner’s map — that is closer to semantic lead time than to a virality score. Outbound prediction is different: it asks what the world can currently hear from a particular claim in a particular form. Hybrid generation is different again: it proposes a new stimulus only when the graph supports a bridge. This article is about the outbound forecast instrument and the selection loop that feeds it. It borrows null-handling and exploration doctrine from publishing-as-sensor work, and leaves the full inbound organ to its siblings.3,6

The paired artefact: Prediction Receipt and Outcome Receipt

This is the deliverable’s core object. Design it completely. Use it before every informative probe — not only the ones you already believe will win.

Prediction Receipt — fields in full

Write this record before publication. Treat it as immutable once the artefact goes live. If you edit the forecast after you see the numbers, you are doing marketing retrospectives, not instrumentation.

Prediction Receipt schema

receipt_id
Stable identifier for this forecast (e.g. pred.2026-07-29.0041).
candidate_id
Pointer to the immutable source exhibit — quote identity, claim grain, or draft stimulus ID. Not a platform post ID (those die).
stimulus_snapshot
Exact text (or hash + retrieved text) of what will be published, plus channel and intended format. Snapshot freezes the object you are forecasting.
semantic_composition
Filled table: concepts/frameworks live; claim and mechanism; tension or contradiction; human implication; evidence type; adjacent IP. Empty cells mean “not ready.”
surface_composition
Hook; length/density; image/layout; structure; CTA; channel and timing. Kept separate so monochrome images are not credited for idea work.
vector_decomposition
Named entries for semantic, authority, temporal, surface and audience vectors — each with the specific priors you are using (e.g. “containment motif rising; carrier unused on this concept; technical audience saturated on credential-theft framings”).
current_conditions
As-of timestamp; concept temperatures and trends in scope; live cases that create doorway context; known fatigue on this quote or framing; exploration-versus-exploitation flag.
expected_outcomes (multi-lane)
One band or qualitative expectation per lane: propagation; audience resonance; conversation quality; action/click intent; market readiness. Optional scheduler scalar, explicitly marked non-learning. No lane may be “whatever happens.”
why (reasoned prior)
Two to five sentences assembling the graph story: which receipts, which edges, which temporal conditions justify the expectations. This is the inspectable selection reason.
uncertainty
What is thin: untested carrier, developing incident details, sparse receipts on this audience slice, high presentation variance historically. Uncertainty is a first-class field, not an apology.
falsifier
A concrete observation that would hurt this forecast. Not “low engagement.” Prefer: “technically informed readers treat the claim as ordinary X rather than mechanism Y,” or “resonance appears only in the prestige-carrier comments, not in mechanism discussion.”
we_will_not_conclude
The sentence the schema requires: explicit non-conclusions. Example: “We will not conclude the concept is false from a null under low exposure; we will not promote packaging folklore from one win.”
exposure_class
Honest sense of who can see this and roughly how much distribution you expect to buy or earn. Without exposure class, null is ambiguous.3
exploration_flag
true if this probe is in the minority exploration budget — internally important but currently cold, or a framing the model dislikes. Exploration slots are first-class in the log, not guilty afterthoughts.3
promotion_path_if_supported
What this run is allowed to feed if lanes behave as expected: candidate hypothesis, packaging prior update, controlled descendant brief — never automatic canon promotion.
created_at / created_by / policy_version
When the receipt was sealed; who or what system sealed it; which forecast-policy or model version produced the prior. Required for honest backtests later.

Outcome Receipt — fields in full

The Semantic Experiment Graph already separates immutable observation from revisable interpretation.2 Keep that split. The Outcome Receipt joins the Prediction Receipt by ID; it does not overwrite it.

Outcome Receipt schema

outcome_id / prediction_receipt_id
Join key. One prediction may eventually have multiple outcome windows (24h, 7d, settled); each window is a row, not an edit of the forecast.
published_at / channel / audience_slice
When and where it actually went live; which slice was addressable. Drift from the prediction’s intended channel is itself data.
exposure_observed
Impressions, reach, or honest proxy. If unknown, say unknown — do not invent a denominator.
outcomes_by_type
Measured or qualitative bands for the same lanes named in the prediction: propagation, resonance, conversation quality, action intent, and any other predeclared lanes. File nulls with the same schema. If you only file heat when numbers spike, you are sampling on the dependent variable.2
relative_band
Against the practitioner’s own baseline, not a fantasy industry benchmark: cool / warm / hot / null, per lane where possible.
prediction_error (typed)
Lane-by-lane: expected versus observed. Examples: “expected warm propagation, received null”; “expected technical resonance, received executive concern”; “expected low click-through, received high saves.” Direction and magnitude matter; a single signed error number does not.
confounders
Platform outages, concurrent newsjack by others, algorithm oddities, posting mistakes, dual-post collisions. List candidates; do not use confounders as a universal escape hatch.
ranked_candidate_explanations
Ordered hypotheses for the error. The Monday-meeting question is not “what won?” but “which explanation did we hurt?”4
priors_to_credit_or_debit
Named only: semantic prior X strengthened; surface prior Y weakened; temporal readiness for motif Z raised; audience slice A less ready than thought. No anonymous “the model should be more confident.”
truth_lane_status
Explicit: unchanged unless independent evidence (not engagement) warrants a separate review. Response must not silently edit this field.
next_falsifying_test
The controlled descendant or follow-up probe this outcome nominates. If the brief cannot name what it is trying to hurt, it is not an experiment. It is content.4
promotion_decision
Human-gated: archive / packaging candidate / hypothesis / hold for isolation / exploration continued. Skipping ladder steps is how folklore gets a logo.2
interpretation_revisable
Observation block sealed; interpretation block may be revised later with a dated amendment — never by erasing the first read.

Illustrative pair (not a measured result)

The following worked pair shows field discipline. It is illustrative, not a measured result. Do not read it as a measured win. Outcome lanes are left as the shape of a join, not filled with fabricated numbers.

Prediction Receipt — illustrative

receipt_id:        pred.illustrative.001
candidate_id:      quote.containment.doorway.12
stimulus_snapshot: short professional-network post;
                   live incident as doorway; no generic
                   hacking language; consequence close
semantic_composition:
  concepts:        agent containment; authority escape;
                   capability acquisition; architecture-not-vibes
  claim:           a demonstrated autonomous boundary failure
                   changes the class of risk, not merely the
                   volume of security news
  tension:         theoretical escape stories vs operational proof
surface_composition:
  hook:            incident doorway, not celebrity
  density:         short explanatory
  structure:       mechanism then implication
  cta:             none — recognition close
vector_decomposition:
  semantic:   containment motif with explicit architecture claim
  authority:  first-party / lab attribution as confirmer class
  temporal:   agent-security rising; fatigue low on this framing
  surface:    incident doorway untested on this exact claim
  audience:   technical and AI-governance readers first
current_conditions:
  as_of: recent multi-party incident week
  exploration_flag: false
expected_outcomes:
  propagation:          warm
  audience_resonance:   hot among technical/AI-governance
  conversation_quality: high (mechanism discussion)
  action_intent:        medium
  market_readiness:     rising
  truth_confidence:     (not a response lane — held separate)
why:
  three independent recent cases converge on the same
  containment mechanism; concept under-rendered relative
  to external temperature; doorway reduces recognition cost
uncertainty:
  incident details still developing; attribution could shift;
  carrier may over-dominate mechanism
falsifier:
  technically informed readers treat it as ordinary credential
  theft rather than autonomous boundary escape
we_will_not_conclude:
  we will not conclude the containment claim is false from a
  null under low exposure; we will not promote “incident
  doorways always work” from one warm result
exposure_class:
  organic professional-network distribution; modest baseline
promotion_path_if_supported:
  packaging prior for incident doorways on containment claims;
  candidate for controlled carrier-swap descendant
created_at: [before publish]
policy_version: forecast-policy.illustrative

Outcome Receipt — illustrative join shape only

outcome_id:             out.illustrative.001
prediction_receipt_id:  pred.illustrative.001
published_at:           [after prediction sealed]
channel:                professional social network
outcomes_by_type:       [NOT FILLED — no fabricated numbers]
prediction_error:       [to be written lane-by-lane after settle]
example_error_shapes_only (hypothetical language, not data):
  - expected warm propagation, observed null
  - expected technical interest, observed executive concern
  - expected lab brand as carrier, observed sandbox-escape
    language driving discussion
truth_lane_status:      unchanged
next_falsifying_test:   same claim, non-incident carrier
                        (hold concept, change carrier)
promotion_decision:     [human-gated after observation]

The learning lives in the typed error, not in a revised forecast. Expected warm, received hot; expected technical interest, received executive concern; expected low click-through, received high saves — each of those sentences nominates a different prior to update.

How a candidate is selected against typed edges

A companion argument already establishes why a cold idea may be a statement about context rather than quality, and why retrieving a position that predates an event differs from manufacturing a hot take. That is the archive-as-option-portfolio case in The Semantic Market Model; this piece cites it and does not re-argue it.1 What belongs here is the operating mechanism: how a candidate option is selected against typed graph edges, and when a hybrid may be generated.

Start from the reverse of ordinary content marketing. Ordinary marketing sees a trend and invents a take. The instrumented path asks: given what is warming in the world and what is already compiled in the canon, which existing claim-grain exhibits have just become timely? Selection is retrieval under constraints, not brainstorming with a heat dashboard.

live conditions (cases, concept temperatures, fatigue)
        ↓
walk typed edges among concept / claim / mechanism nodes
        ↓
retrieve independently receipted exhibits (quotes, claim grains)
        ↓
score candidates as reasoned priors (multi-lane), not one scalar
        ↓
seal Prediction Receipt (incl. falsifier + exposure class)
        ↓
publish bounded probe
        ↓
join Outcome Receipt → typed error → named prior updates

In a running system the grounding step is not “ask a model what this reminds you of.” It is closer to: take the candidate’s text and article context; search the compiled concept graph with the candidate as query material; accept only references that resolve to real page identities; keep unmatched strings as observations (negative space that may indicate a missing concept), never as invented edges that mint authority. Heat that later files against concepts is advisory — resonance, fatigue, temperature — and must not reorder truth authority in the canon. That separation is the mechanical expression of the governance rule this article ends on.

Walk a concrete selection, not a slogan. Suppose a live incident week has raised temperature on agent containment and on authority-boundary failure, while an older claim-grain exhibit in the archive already argues that architecture, not vibes, is the durable response to escape stories. The system does not start by drafting a new hot take. It asks which existing exhibits already carry the mechanism; which typed edges connect those exhibits to the warming concepts; whether the exhibit is under-rendered or fatigued; and whether the incident is being used as a doorway to a pre-existing claim or as an excuse to invent shallow novelty. Only then does it generate the actual post surface for judgment — because evaluating what a quote becomes as a post is stronger than scoring an isolated sentence — and only then does it seal the Prediction Receipt. If the graph cannot name the bridge, the candidate is rejected or demoted to pure exploration with the flag set honestly.

Notice what is still missing in many operating stacks even when grounding and outcome heat already exist: the sealed multi-lane prior itself. Ingredients for learning — relative engagement baselines, null outcomes, concept temperature, repeated post episodes, caption and card context, exploration doctrine — can be present while a formal before-publication model still is not. The Prediction Receipt is the missing artefact that turns those ingredients into an error-producing instrument. Without it, the system can reprice and schedule; it cannot yet say “expected warm resonance among technical readers, uncertainty high because this carrier is untested, falsifier X” and mean it as a ledger entry.

Selection criteria that stay legible:

  1. Edge support. Prefer candidates whose semantic composition is already linked by typed relations (extends, contradicts, mechanism-of, evidence-for) rather than keyword overlap with a trending tag.
  2. Receipt history. Prefer exhibits with prior outcome receipts you can condition on — including nulls — over shiny new wording with no history.
  3. Temporal doorway. Prefer a pre-existing claim that a live case makes newly hearable over a fresh sentence that merely mentions the case.
  4. Under-rendered versus fatigued. High external temperature plus low own rendering is a different prior from high own fatigue; do not average them into “post more of X.”
  5. Exploration budget. Reserve a minority of publishes for cold internal priorities the model would not pick. A sensor that only samples where it expects signal becomes a mirror with a content calendar.3

The evaluator path that content systems already approximate — generate the actual post surface, then judge — is the right order for surface composition. The Prediction Receipt adds the graph comparison before the final go decision: decompose, compare with the learned heat graph, seal the multi-lane prior, then publish. Judging an isolated quote is weaker than judging what the quote becomes as a post; judging a post without a sealed prior is weaker still, because you cannot form prediction error.

Cross-Heat Hybrid — the gate, not the soup

Cross-Heat Hybrid is an established pattern: begin from independently hot concept nodes, find a shared or bridging neighbour, and propose a stimulus that makes the bridge explicit. It is graph-guided only when the heat traces back to real receipts. Otherwise it is brainstorming with better vocabulary.

The operating rule for this piece is narrow and enforceable:

Hybrid generation rule Generate a cross-heat hybrid only when typed graph edges support a real bridge between independently receipted concepts. Independently hot is not enough. Shared tag vocabulary is not enough. A model’s aesthetic preference for combining fashionable words is not enough. If you cannot name the bridge edge (shared mechanism, contradiction, empty-space neighbour already present in the graph) and the supporting exhibits, do not mint the hybrid.

The disciplined sequence:

1. Find independently hot concepts (receipt-backed temperature).
2. Walk their typed edges.
3. Identify a shared mechanism, contradiction or empty-space neighbour.
4. Find existing articles and quotes that already support that bridge.
5. Draft one claim-grain probe (not five-tag jargon soup).
6. State why the combination was nominated (edge IDs / page IDs).
7. Preregister multi-lane expectations and a falsifier.
8. Publish and obtain a new outcome receipt.

That is heat-guided synthesis. The failure mode it blocks is prompting “combine the five hottest concepts into a post,” which teaches the system to chase itself and produces semantic soup that no outcome can address at concept grain.

A timely-quote retrieval episode has the same discipline with less novelty risk: an old exhibit whose concept becomes timely after a live incident is exercised as an option on prior work, not rewritten into a shallow reaction. The forecast still names which vector — usually temporal doorway plus semantic composition — is carrying the bet, so a miss can debit the right prior.

What the error is allowed to change

Prediction error is the point of the instrument. The danger is not that forecasts will be wrong; they will be. The danger is that a compounding system becomes excellent at repeating what already receives attention, and quietly lets the market author the canon.

Write error back into named priors only:

Error pattern Prior it may update Prior it must not update
Expected warm resonance, received null, exposure honest Market readiness; framing effectiveness; audience fit; fatigue Truth of the claim; canon membership
Expected technical conversation, received broad shallow propagation Surface/hook priors; doorway priors; audience-slice targeting “The idea is validated because it travelled”
Wrong-direction forecast (expected cool, received hot) Underestimated readiness; missing higher-level concept; doorway power; or engagement trap flags Automatic promotion up the canon ladder
Same claim, different renderings, high variance Presentation-effect band; surface priors Idea quality from a single lucky card
Same claim, repeated renderings, shared response shape Idea-effect candidate (still provisional) Final truth determination from audience alone

Repeated renderings of the same claim are how you separate idea effect from presentation effect without pretending one post can do both jobs. Variance across renderings is a presentation-effect band; shared response shape is an idea-effect candidate.4 Distribution effects — timing, platform reach, audience conditions — need their own accounting, or you will credit the concept for a lucky algorithmic hour.

Controlled descendants remain the right next-step language when an outcome nominates a follow-up: hold concept change carrier; hold carrier change concept; hold structure vary both. Practical order often starts with carrier swap when celebrity or incident doorways are loudest.4 The Prediction Receipt for the descendant should name which explanation the new probe is trying to hurt.

Hard boundary The audience may teach the system how an idea travels. It may not decide whether the idea is true. Heat may adjust readiness and packaging, not truth. Truth, propagation, market readiness and audience resonance remain separate lanes. Canon promotion remains human-gated. Observations and interpretation remain distinct.

That boundary is not etiquette. It is what separates a learning worldview from a sophisticated engagement addiction machine. Publish only what the model predicts will heat, and you get the second failure mode that looks like success: a calendar that mirrors last week’s rewards while the exploration budget starves.3

Internal significance versus external heat should remain a matrix, not a blended score. The off-diagonal cells are diagnostic: high internal significance and low external heat may mean dormant, badly framed, or early; low internal significance and high external heat may mean a missing higher-level concept, an underestimated market problem, a powerful doorway, or a shallow trap. All four findings are useful. Averaging them into one “priority” number destroys the diagnosis.

A companion piece will take up the relational ontology of heat — what heat is as a multi-party relation rather than a post property. This article only needs the operational consequence: if heat is relational, your receipts must keep the relation’s parts nameable, and your write-back rules must forbid the market from editing the truth lane.

Practically, error write-back should look boring and procedural. A weekly review does not ask “what won?” It samples alerts you would have predicted as warm, nulls you predicted as informative, wrong-direction surprises, and exploration slots. For each joined pair it names one prior update in a sentence a skeptic could audit: “Debit readiness for scarcity-migration framings with this audience under mid-week organic distribution,” not “the algorithm hated us.” Policy changes that follow from many such sentences — fatigue half-lives, doorway priors, exploration share — should be versioned and, where possible, challenged by replaying historical freezes rather than silently rewritten by an evaluator agent optimising against its own prose. AI may propose the defect; deterministic comparison against frozen cases should generate the evidence; the practitioner remains the policy owner.

How to know the instrument works (proposed, not yet run)

Doctrine without instruments rots into lore. The minimum proof burden for Prediction Receipts is real, and honesty requires stating what has not been run.

Proposed — not yet run The following protocols are specified so a reader could execute them next week on their own corpus. They are not reported results. No historical backtest score, no live predicted-versus-actual table, and no idea-mean versus surface-variance estimate is claimed in this article.

1. Ten historical backtests without future leakage

Question. Do multi-lane Prediction Receipts, sealed against frozen historical state, produce better-calibrated and more reusable errors than a scalar engagement predictor?

Method (runnable protocol).

  1. Build a settled episode set. Select at least ten past publication episodes that have: immutable stimulus text; channel and approximate exposure class; outcome observations by type (even if some lanes are qualitative); concept grounding that existed at the time or can be reconstructed from a versioned graph. Prefer variety: wins, nulls, and mixed lanes.
  2. Freeze the state tuple for each episode. For each target publish timestamp T, restore or reconstruct: wiki/graph commit as-of T−ε; concept temperature and fatigue as-of T−ε; policy/model version identifiers; any live-case context that was actually available before publish. The critical protection is the future-leakage rule: the forecast machinery must not see a graph or temperature field that already knows how the episode ended. If gold already contains the settled interpretation of the week’s outcomes, tag the row leakage and report it separately — do not average clairvoyance into skill.
  3. Seal retrospective Prediction Receipts under freeze. Using only as-of-T−ε state, fill the full schema: compositions, vectors, multi-lane expectations, uncertainty, falsifier, we-will-not-conclude. Human or model may draft; a second reviewer checks for hindsight language (“because it went viral…”) and rejects contaminated seals.
  4. Join actual outcomes without editing the forecast. Score lane-by-lane calibration (ordered bands), falsifier hit rate, and whether ranked explanations from the original error write-up were specific enough to name a prior. Compare against a scalar baseline (e.g. single predicted engagement band) on the same freezes.
  5. Report distributions, not a hero number. Include null episodes. Split by exploration versus exploitation if flags exist. Publish contamination rate (how many rows failed the leakage audit).

What would count as support. Multi-lane receipts yield more actionable prior updates per episode than scalar baselines; falsifiers are empirically usable; leakage-tagged rows are few and do not drive the headline.

What would falsify. After honest freezes, typed receipts add no decision value; or performance appears only when future state leaks into the prior.

Status. Proposed. Not run as a published study here.

2. Three live predicted-versus-actual pairs

Question. Does the instrument work in forward time, including when the market is quiet or moves the wrong way?

Method.

  1. Pre-register three probes with full Prediction Receipts before any publish.
  2. Include by design: at least one exploration-flagged or low-expected probe (null-seeking is allowed); at least one probe where a wrong-direction outcome is thinkable; at least one exploitation probe on a warm territory.
  3. Publish; do not edit receipts.
  4. Settle outcomes on a predeclared window; file Outcome Receipts with typed errors and human-gated promotion decisions.
  5. Record whether truth_lane_status remained unchanged absent independent evidence.

What would count as support. All three produce interpretable typed errors; at least one null is filed without narrative collapse; no automatic canon write from engagement.

What would falsify. Receipts unused in decisions; outcomes rewritten into the prediction; truth lane quietly updated from likes.

Status. Proposed. Not run as a published live set here.

3. Repeated-rendering estimate of idea-mean versus surface-variance

Question. For one claim-grain exhibit, how much of observed response variance is presentation versus shared idea shape?

Method. Hold the claim constant; publish or queue multiple surface compositions (carrier swap first if doorway effects are loud). Seal a Prediction Receipt per rendering. After settle, estimate: shared response shape across renderings (idea-effect candidate) versus variance band (presentation-effect). Do not promote a single lucky rendering into “the idea works.”

Status. Proposed. Not run as a published estimate here. The separation doctrine is already established; the numerical estimate for any particular corpus is not.4

These three instruments are deliberately uneven in cost. The live triple can start on the next three publishes. The repeated-rendering study needs one claim you are willing to re-surface. The ten-episode backtest is the heavy lift, and it is also the only way to stop a beautiful schema from becoming untested folklore. Until it runs, the correct public language is the language this article has used: the instrument is designed; the learning loop is specified; the measured calibration is not yet claimed.

What you can do on Monday without waiting for the study

You do not need ten backtests completed to stop learning nothing from your own publishes.

Sibling boundaries, briefly. The whole-object market model is 193. The three clocks of a learning system are 194; semantic lead time is 195 — inbound timing, not outbound forecast receipts.5,6 Publishing as an active sensor already argued that concept-sized probes return outcome evidence; this extender adds the sealed before-publication forecast and the error loop.3 The relational ontology of heat is forthcoming and is not developed here.

The scarce asset is an inspectable reason for selecting one variant — and a record of what the miss taught.

Generative systems will keep making variants cheaper. Prediction Receipts are how you stop that abundance from becoming either superstition or addiction. Decompose the candidate. Seal multi-lane expectations and a falsifier. Publish. Join the outcome. Update readiness and packaging. Leave truth alone. If the forecast was wrong in a nameable way, you learned. If it was right, you still learned only what the lanes support — and nothing more.

References

  1. Scott Farrell / LeverageAI. “The Semantic Market Model.” — Whole-object market model; archive-as-option-portfolio argument cited, not retaught. https://leverageai.com.au/wp-content/media/articles/article.php?article=193-the-semantic-market-model
  2. Scott Farrell / LeverageAI. “Semantic Experiment Graph” (Two Compositions and the Outcome Receipt, ch3, cite key #9e3101). — Two-composition preregistration; outcome receipt as observation plus interpretation; non-interchangeable outcome types; null filing; promotion ladder; “we will not conclude X.” https://leverageai.com.au/wp-content/media/articles/157-semantic-experiment-graph.html
  3. Scott Farrell / LeverageAI. “Publishing Is an Active Sensor” (Null response and the exploration budget, ch8, cite key #bd018c). — Null is a measurement with expectation and exposure class; exploration budget first-class; sensor must not only sample where it expects signal. https://leverageai.com.au/wp-content/media/articles/158-publishing-is-an-active-sensor.html
  4. Scott Farrell / LeverageAI. “Semantic Experiment Graph” (Controlled Descendants: Carrier, Concept, Structure, ch6, cite key #79c348). — Controlled descendants; idea-effect versus presentation variance across repeated renderings; “which explanation did we hurt?”; illustrative-not-measured labelling convention. https://leverageai.com.au/wp-content/media/articles/157-semantic-experiment-graph.html
  5. Scott Farrell / LeverageAI. “The Three Clocks of a Learning System.” — Bronze, queue and gold on unequal clocks; named as substrate, not retaught. https://leverageai.com.au/wp-content/media/articles/article.php?article=194-three-clocks-of-a-learning-system
  6. Scott Farrell / LeverageAI. “Semantic Lead Time.” — Inbound motif completion before market attention; distinct from outbound Prediction Receipts. https://leverageai.com.au/wp-content/media/articles/article.php?article=195-semantic-lead-time