Leverage AI

Attention Flight Recorder: Quiet Needs a Decision Log

A quiet intelligence system is trustworthy only when every interruption and every suppression is reconstructable and replayable — because low alert volume cannot tell excellent negative cognition from a broken sensor.

Scott Farrell · LeverageAI · Long-form article

Two systems can look identical on a dashboard.

A sophisticated system saw thousands of things and correctly suppressed almost all of them.

A broken system failed to see anything.

Both produce low alert volume. Both can produce a human who says “it was a quiet week.” Only one of them did real work. The difference is not how many times the system spoke. The difference is whether every interrupt — and every deliberate silence — can be reconstructed, sampled, and replayed against an alternative policy without hindsight leaking into the test.

Quiet is not evidence of intelligence. Quiet with a reconstructable decision log is.

That is the claim of this piece. Personalised attention systems are finally good at being quiet. Personalisation is the product: a worldview map of what the system owner cares about, which distinctions matter, which concepts connect, which sources carry weight. But personalisation is also the principal overfitting risk. Learning only from what the radar selected turns an echo chamber into code. The moat work already names the guards this problem requires: an alien-signal lane, a suppression audit, and a replayable trace. Missing guards make compounding unsafe. This article specifies the suppression audit in full, and the decision artefact that makes the other two guards operational rather than aspirational.

Honesty constraint This piece specifies an instrument. It does not report measurements from that instrument at scale. There is no populated week of the full decision-record schema presented as a completed experiment; no human-labelled set across the four review buckets; no baseline-versus-two-candidate policy replay that has been executed and scored for this article; no assembled regression corpus of late bloomers, resurrections, and correct long silences claimed as already built here. Canon describes the shape such a corpus should take; that shape is cited, not claimed as a finished asset. Qualitative practice from a running personal intelligence radar — decision logs retained, near-alert review imagined, a multi-pass suppress-then-push trace — is legitimate to describe as existing shape. Invented alert counts, precision/recall figures, and “tuned” half-life results are not. Keep the line: instrument fully specified; nothing measured for this piece.

Where this sits — and what it refuses to re-teach

This is the eighth article in a nine-piece run about a semantic-market learning system. Seven siblings are live. They already own the market object, the three clocks of memory growth, semantic lead time, prediction receipts, heat as a relationship, case formation, and the judgment-join control plane. This piece does not re-argue those axes. It consumes them as the world the attention system operates on.

Two prior frameworks are load-bearing and must stay in their own jobs.

Reflexive Agent Design already owns path inspection and the honesty device of labelling which 2×2 cell a walk fell into — including a queue-specific translation: correct reasons, right case, appropriate canonicality, right time. That chapter scores walks. This piece records decisions, including the silent ones, and specifies which decisions to pull when most decisions are silences.

Replay-Driven Design Evolution already owns the freeze protocol, the Future-Leakage Rule, and the exact replay pin: wiki commit + queue checkpoint + event-log offset + agent/tool configuration. Replay without a freeze protocol is cosplay science. This piece does not rewrite that protocol as if it were new. It places attention-policy candidates on top of it.

A later piece in this run will own latent questions — the way a persistent system can close a question the human never typed. That article is not published yet. It is named here only so the boundary is clear: this piece owns the decision record and the sampling/replay loop that calibrates attention policy, not the latent-question phenomenon itself.

The reader question is practical: how do you tune a personalised attention system without making it noisy, self-confirming, or silently blind? After this piece you should be able to retain an Attention Flight Recorder, sample four review buckets that catch different blindnesses, and compare policy changes against the same frozen past — multi-axis, human-gated, without future leakage.

Two loops, three things learned

The first loop is already familiar from the rest of the set:

world
→ observations
→ cases
→ wiki-grounded interpretation
→ attention decision

That loop interprets the world for the system owner. The second loop is what turns a personalised feed into a self-calibrating intelligence system:

attention decisions
→ retained walks and decision history
→ AI review
→ policy diagnosis
→ freeze-safe replay of alternatives
→ human-gated policy revision

The second loop evaluates the interpreter. Three different objects are being learned, at different speeds, and they must not be collapsed into one “AI tuning” story.

The worldview is the slower gold layer: what the owner cares about, which distinctions matter, which concepts connect, which people and institutions are load-bearing in which domains, which prior episodes have been confirmed or weakened. Worldview change is rare and deliberate.

The attention policy is the queue machinery: how quickly heat rises and fades, how often a case is revisited, how relevance affects cadence, what crosses an alert threshold, when a quiet case expires, when new evidence reactivates it. These constants start as informed guesses. They need not remain folklore.

The cognitive apparatus is how the model reaches a judgement: which wiki pages it opens, whether it reaches canonical material, whether it thrashes through broad searches, whether attribution walks terminate too early, whether the tool surface elicits bad behaviour. Retained walks expose this layer. Path quality and answer quality are different questions — that is Reflexive Agent Design’s ground — and a correct decision through a bad path remains dangerous because it may only have been lucky prior.

Cadence parameters — half-lives, multipliers, revisit windows — appear in this piece only as policy knobs under test in a replay. The growth-rate argument for why those clocks exist at all belongs to the three-clocks sibling. Do not confuse a knob you can A/B with a doctrine you re-derive.

The attention-decision record

Operational telemetry — scrapes, token counts, durations, failures — is necessary and insufficient. Counts tell you volume. They do not tell you why this case received the owner’s attention while that one remained silent. The primary artefact is not a dashboard. It is a decision record that makes negative cognition as inspectable as a push notification.

Call the full stack an Attention Flight Recorder: not only what the system did, but enough structure to reconstruct why. Five related evidence classes usually cohabit in a real system; the flight recorder is the join of them for attention work:

Record class What it captures
Operational telemetry Scrapes, scans, model calls, tokens, durations, failures
Decision ledger Attach, create, merge, reprice, fade, surface, suppress, resolve, reopen
Cognitive provenance Which wiki pages, relationships, and bronze sources the agent actually walked
Case trajectory Every state the case passed through, with evidence and interpretation at each point
Outcome / feedback What later happened; whether the owner agreed, ignored, upvoted, or downvoted

What no prior chapter in this set specifies field-by-field is the single attention-decision record: one interrupt-or-suppress decision, retained with the same seriousness whether the human was interrupted or not. The delta from path scoring is simple and load-bearing: suppressions get recorded exactly as carefully as alerts.

Specify the fields. A reader should be able to implement them next week.

Field 1 — Trigger evidence What observation or re-observation caused this evaluation? Include source identifiers, evidence class (press retelling, first-party reconstruction, engagement spike, null reobservation), timestamps, and the cheap signals that nominated the case for judgement. Without the trigger, every decision looks like a free-floating opinion.
Field 2 — Case state at decision time Snapshot the case as the decision saw it: identity, heat band, uncertainty, relevance to the owner, open questions, prior disposition history, and the evidence set then attached. Case formation and the judgment join own how that state is built and mutated; the flight recorder freezes a readable cut of it at decision time so later review is not reconstructed from memory.
Field 3 — Wiki walk (cognitive provenance) Which pages, edges, and bronze objects were actually opened? Did the walk reach a canonical anchor before synthesis? Which branches were abandoned? Path length, redundant reads, and explicit cell classification belong here — that is the walk scorecard Reflexive Agent Design already requires. The decision record points at the walk; it does not replace the scorecard.
Field 4 — The decision Interrupt or suppress (and the finer verb if the ledger is richer: hold, fade, expire, reprice-without-push, push, reopen). Include the structured reason codes the control plane can store — not only free prose. Prose reasons are easy to invent after the fact; structured codes are harder to fake and easier to aggregate.
Field 5 — Policy and model versions in effect Heat formula version, half-life and cadence constants, threshold and hysteresis settings, prompt/tool/model identifiers for any AI adjudication, and configuration hashes. A decision without versions cannot be replayed. A “we made it quieter” story without a versioned policy is folklore.
Field 6 — Mutation caused What durable state change, if any, did this decision apply? Queue position, heat write, suppress flag, push payload, alias, expiry. The judgment join already insists mutation is not narration; the flight recorder stores the mutation receipt next to the attention decision so the two cannot drift.
Field 7 — Later outcome (backfilled) The field most systems never implement. Once the world has moved on: did a suppressed case turn out to matter? Did an alert turn out to be noise? Did first-party confirmation arrive after the review window that held the case quiet? Did the owner act, ignore, or disagree? Outcome is not available at decision time. It must be a first-class nullable field that later process fills — human label, delayed evidence join, or both. Without backfill, review is only aesthetics; with it, negative cognition becomes a supervised problem.

The minimum viable flight recorder is not “we log prompts.” It is these seven fields, written for every attention decision, including the ones that produced silence. If suppressions are cheap to omit, they will be omitted — and the training distribution collapses to the positive surface.

A qualitative receipt: silence that earned its keep, then an interrupt that did too

A concrete production-shaped trace from a personal intelligence radar shows why the record must retain the quiet passes, not only the push. A case about a public agent-security incident sat open. Across several re-observation windows the system attached nothing new: null reobservation, causal attribution still stalled, review held event-driven rather than engagement-driven. Those suppressions were not absences. They were repeated judgements that the evidence object had not changed.

Then new material arrived: a first-party technical reconstruction on the victim side, and independent reporting that broadened the containment picture beyond a single organisation. The case was marked dirty, evidence attached rather than duplicated into a new case, state repriced from stalled press retelling toward artifact-backed analysis, and a push issued because the implication for the owner’s active containment work had crossed the interrupt line.

Handle the public disclosure stack the way other pieces in this run already do. Hugging Face’s first-party incident write-up is directly fetchable.1 OpenAI’s own disclosure page is treated as unreachable for citation purposes in this writing environment (HTTP 403 is the house pattern across siblings); OpenAI-attributed claims ride independent reporting instead.2 Independent reconstruction of the public document stack is also available.3

What matters for this article is not the incident’s full causal story — lead time and case formation own deeper cuts. What matters is the decision shape:

repeated null reobservations → suppress (justified silence)
new first-party + corroborating evidence → attach (not duplicate)
evidence-class and structural change recognised → reprice
implication for owner’s active work material → push

Without the suppress rows, the push looks like magic or like a chatty monitor that finally got lucky. With the suppress rows, you can see negative cognition doing work: the system tolerated silence, maintained the case, refused to re-alert on null traffic, and spent the interrupt only when the evidence object changed class. That is the central trust claim a flight recorder exists to support — and it is a qualitative claim from a trace shape, not a measured precision figure.

Four review buckets — the suppression audit specified

The weekly artefact is not a general activity dashboard. It is an attention review: a decision book. Near misses alone are valuable and insufficient. Near-miss review is boundary sampling — cases close enough that a small policy change would have flipped the outcome. That teaches calibration of the existing scoring geometry. It does not reveal news that sat miles below the threshold because the worldview lacked the vocabulary to recognise it.

Four buckets are required. Each catches a different blindness. Together they are the suppression audit the moat chapter requires but does not specify.

Bucket 1 — Surfaced alerts

These test precision: did the system interrupt the owner for something genuinely worthwhile? For each alert in the sample, the review asks whether the owner agreed at the time, whether later evidence confirmed consequence, and whether the interrupt was early enough to matter. Alerts are the easy bucket. Everyone already looks at what fired. The flight recorder’s job here is not discovery of the alerts — it is retaining enough structure that “I liked it” is not the only label available six weeks later.

Grade the walk under the queue translation already published: correct reasons (independent corroboration / first-party confirmation / trajectory — or a single noisy aggregator spiking?), right case (joined the existing case or spawned a duplicate?), appropriate canonicality (provisional anchor promoted when a primary arrived, without deleting bronze?), right time (watch / alert / resolve windows matching urgency, or FIFO sludge?). Output-only success — “it showed up somewhere” — is the queue’s version of the lucky answer.

Bucket 2 — Near misses

These test decision-boundary calibration: were cases just below the line correctly held back? The practical question the system owner already asks — “what nearly got to alerting status this week?” — belongs here. For each near miss, record the exact factor that held it below the line: heat still short, confirmation missing, relevance too low, review window expired before first-party arrival, wrong wiki entry point, duplicate suspicion. Near misses are where small constant changes would flip outcomes. They are also where it is tempting to “fix noise” by raising thresholds until the boundary sample disappears. That is optimising for fewer alerts, not for better judgement.

Bucket 3 — Random (and confident) suppressions

These test recall — and they are the non-obvious bucket. Near misses sit close to the line. Confident suppressions sit far from it. Sampling only the boundary never asks: what did the system dismiss with high confidence that later proved important?

Personalisation makes this bucket mandatory. The radar is intentionally shaped by the owner’s map. That is its power. It is also how a blind spot becomes structural. A case can be suppressed not because evidence was weak, but because the worldview had no place to hang the significance — “already known” misread as “not consequential,” a domain the gold layer underweights, a relationship the walk never enters. Those failures do not appear as near misses. They appear as silence with a clean conscience.

Sampling design matters. Pull a stratified sample of suppressed decisions, not only the ones an evaluator agent finds narratively interesting after the fact. Include cases the policy marked low-relevance with high confidence. Include expired quiet cases. Where later outcome is known — a “late bloomer,” a resurrection, an external event that made the quiet case obviously material — backfill Field 7 and treat the row as gold for policy diagnosis. Where outcome is still unknown, label that honestly rather than inventing a retrospective story.

This is the heart of the suppression audit: force the training distribution to include the negative surface, especially the confident negative surface. Learning only from what the radar selected is how an echo chamber becomes code.

Bucket 4 — Alien or magnitude cases

These test worldview closure: what looked significant in the world but had no detectable intersection with the owner’s existing map? Magnitude without intersection is a different failure from “we scored it and suppressed it.” The system may never have formed a case, or may have formed a thin one that died without a walk into gold. The alien-signal lane is the sibling guard: a small allocation that forces inspection of high-magnitude, low-intersection events so the map can be challenged rather than only reinforced.

Bucket 4 is where you catch the failure mode near-miss review cannot: the scoring geometry never engaged. If you only improve precision and boundary calibration, you polish a local optimum. Alien sampling is how personalisation stays a feature instead of becoming a sealed room.

Bucket Primary question Blindness it catches
Alerts Was the interrupt worth it? Precision failure; lucky paths that look like skill
Near misses Was the boundary right? Threshold / heat / timing miscalibration
Confident suppressions What did we dismiss that mattered? Worldview-shaped recall failure; clean-conscience silence
Alien / magnitude What mattered with no map intersection? Worldview closure; never-scored significance

AI is particularly good at reading these trajectories once the records exist. Deterministic analytics can count near-threshold cases and mean lifetimes. Useful — and still counts. An analyst model can say: these cases faded before first-party confirmation commonly arrives; these suppressions repeatedly treated “already known” as “not consequential”; attribution corrections only fire after a secondary is fetched; the walk keeps entering through broad concepts and missing the owner’s containment frameworks. That is decision-book review, not a dashboard. The trustworthy material remains the stored evidence, policy version, wiki neighbourhood, structured decision, and mutation — not a persuasive retrospective paragraph generated afterward.

The objective is a trade-off, not a volume target

Being quiet is valuable. Being quiet forever is not. Being noisy is not either. The quality target is not “few alerts.” It is:

Maximum justified silence with timely interruption on the few things that matter.

Those are two objectives held at once. Maximise silence that can still show its working. Do not sacrifice early interrupt on the material minority. Trust, not throughput, is the posture: write the shadow-run checklist before you optimise for a quieter week.

Trap A — Optimise for fewer alerts Raising thresholds, accelerating fade, and suppressing near misses will make the dashboard look sophisticated. It also makes a broken sensor look the same. Quiet without reconstructable judgement is muteness. If your primary success metric is alert count, you will train the system to hide.
Trap B — Optimise for higher approval If the system only surfaces what the owner already likes, approval rates climb and the map never gets challenged. Rubber-stamp evaluation is not precision; it is the echo chamber with a thumbs-up. Approval is a useful label inside Bucket 1. It is a disastrous sole objective.

The dual objective forces multi-axis reading of every proposed policy change: did justified silence increase without late misses? Did interrupts arrive earlier on the cases that deserved them? What did cognition cost? What happened to alien-lane coverage? Those questions cannot collapse into one number without losing the plot.

The single-score warning

When you compare a baseline attention policy to a candidate — a longer heat half-life, a stronger first-party confirmation prior, a different cadence multiplier — the temptation is to report one aggregate: “candidate wins.” That single score is how regressions hide.

A candidate can improve average approval while missing a late bloomer class. It can cut cognition cost by skipping walks that would have reached canonical sources. It can look better on precision because it alerts later — after the world is already obvious — and thereby destroy lead time. It can win on a blended score while losing on the exact cases the regression corpus exists to protect.

Reflexive Agent Design’s line for walks is that the cell label is the honesty device: without it, teams optimise path length into silence or accuracy into luck. The attention analogue is multi-axis comparison with explicit cells, not a single quality float:

If a policy “wins” only by hiding a cost, recall, or lead-time regression behind an average, it has not won. Refuse promotion on a single score the way you refuse a green answer without a path label.

Counterfactual attention calibration under freeze

Because bronze observations, serial case history, decision records, wiki walks, and policy versions are retained, alternative attention policies can eventually be replayed against the same past. That is stronger than changing a constant and hoping next week is informative.

Example shape — constants are illustrative policy knobs under test, not tuned results from this piece:

Baseline:
  heat half-life = H0
  alert requires relevance ≥ medium
  hot cadence multiplier = C0

Candidate A:
  heat half-life = H0 + Δ   (longer memory of heat)

Candidate B:
  first-party confirmation adds a stronger trajectory prior

Candidate C:
  high relevance slows fading but does not directly raise heat

Compare, on the same event history: which cases alerted; how early; which false interrupts appeared; which important cases remained silent; cognition cost; whether attribution was correct at alert time; which wiki paths supported the decisions. Report each axis. Do not average them into a trophy.

The replay substrate is already specified in canon. Restore the git-pinned tuple:

wiki commit
+ queue checkpoint
+ event-log offset
+ agent/tool configuration

Feed subsequent observations serially, not batched. Serial state is the phenomenon under test. Run controlled arrival-order permutations when order itself is a hypothesis: secondary first, primary first, aggregator first, primary delayed, one source missing, engagement burst before technical corroboration. Compare designs on the same corpus.

The protection that makes this science rather than cosplay is the Future-Leakage Rule: do not replay the final period against a wiki that already knows what happened during that period. Otherwise you are not testing discovery. You are testing whether the system can read a spoiler. Future leakage is the silent invalidation mode for historical replay — name it, forbid it. At freeze time, record the as-of timestamp of world knowledge, which events are allowed after it, which variants saw which permutations, where decision traces are stored, and what would count as contamination. If any answer is “the model just knew,” you do not have replay. You have a demo.

The regression corpus need not be the whole world. Canon already describes a useful shape: a modest hand-labelled set of important cases, convincing noise cases, canonical-source-late cases, hot-then-dead cases, and slow-burn cases that initially looked unimportant. That list is the target shape for late bloomers, resurrections, and correct long silences. This piece cites that shape as the design target. It does not claim the corpus has been assembled and used as a completed experiment here.

AI proposes; humans own the policy

AI should inspect histories, name recurring defects, and propose policy changes. It must not silently rewrite heat constants, thresholds, or alert rules from its own reading of its own decisions. That creates a self-modifying system optimising against its own interpretation of its own past — a closed loop with fluent explanations and no external brake.

Failure mode — evaluator agent as silent policy author An agent that both reviews the flight recorder and applies policy mutations collapses challenger and owner. Persuasive retrospectives are cheap. Versioned, freeze-tested, human-approved policy changes are the product. If the evaluator can rewrite the attention policy without a human gate, you no longer have a calibration loop. You have an unsupervised policy drift generator with a changelog written after the fact.

The safer loop is explicit:

AI inspects histories
→ names a recurring defect
→ proposes a policy change (falsifiable, with affected historical cases)
→ deterministic freeze-safe replay compares baseline and candidate
→ human reviews the changed cases across the four buckets
→ approved policy is versioned
→ later outcomes become the next evidence

AI is the challenger and analyst. Replay is the evidence generator. The system owner remains the policy owner. That is decision-system CI/CD for attention allocation: guess → observed decisions → labelled outcomes → diagnosis → counterfactual replay → policy revision → regression corpus. The product under continuous build is not a recommendation engine. It is the owner’s allocation of attention.

What a weekly attention review contains

When the records exist, the weekly review can stop asking “did the radar work?” and start asking a better question: where is the radar systematically allocating attention incorrectly, and what change would improve that without making another part worse?

Useful sections, all grounded in flight-recorder rows rather than vibes:

That review is the primary operating artefact. The flight recorder is the substrate. Freeze-safe replay is the challenge procedure. Human approval is the promotion gate.

What this piece does not claim — instruments still unrun

Architecture articles are often judged on whether they sound finished. This one should be judged on whether it admits what it has not measured and still leaves a reader able to measure it.

Proposed instrument 1 — One populated week of the schema Retain Fields 1–7 for every interrupt and suppress decision for seven consecutive days on a real personal intelligence radar. Success is completeness and reconstructability: can a third party replay why each silence and each push happened without asking the model to remember? Not yet run as a reported experiment for this piece.
Proposed instrument 2 — Human labels across four buckets Sample alerts, near misses, confident suppressions, and alien/magnitude cases. Label owner agreement, later consequence, and path cell where a walk exists. Report each bucket separately. Not yet run for this piece.
Proposed instrument 3 — Baseline versus two candidates on one freeze pin Pin wiki commit, queue checkpoint, event-log offset, and agent/tool configuration. Replay the same serial event window under baseline and at least two policy candidates with no future leakage. Compare multi-axis deltas; refuse single-score promotion. Not yet executed as a scored comparison for this article.
Proposed instrument 4 — Regression corpus of the canon shape Hand-label important cases, convincing noise, canonical-source-late, hot-then-dead, and slow-burn rows; include late bloomers, resurrections, and correct long silences. Use as a promotion gate for policy changes. Target shape is canon; a completed corpus for this piece is not claimed.

Existing practice already includes decision logging and the idea of reviewing cases that nearly alerted. That is real qualitative substrate. It is not a substitute for the four instruments above. Readers should treat the Attention Flight Recorder as a fully specified design with production-shaped receipts — not as a benchmark winner with invented statistics.

What you should be able to do now

Take any personalised monitor, news radar, or institutional attention queue and force the calibration loop out loud:

  1. Do suppressions leave the same record as alerts? If silence is only “nothing written,” you cannot distinguish skill from blindness.
  2. Are the seven fields present? Trigger evidence, case snapshot, wiki walk, decision, policy/model versions, mutation, later outcome backfill.
  3. Which of the four buckets do you actually sample? If the answer is “alerts and maybe near misses,” confident suppressions and alien cases are still free to rot.
  4. What is the objective? Maximum justified silence with timely interruption — not fewer alerts, not higher approval alone.
  5. Can you name a policy change as a falsifiable candidate and pin a freeze tuple? Wiki commit, queue checkpoint, event-log offset, agent/tool configuration — serial replay, Future-Leakage Rule enforced.
  6. Do you report multi-axis deltas or a single score? If a candidate wins only on a blend, demand the cost, recall, and lead-time lines.
  7. Who is allowed to change the policy? If an evaluator agent can rewrite attention constants without a human gate, stop and insert the gate before you optimise anything else.

You do not need a year of labels to start. Start retaining the decision record this week, including the quiet rows. Sample one confident suppression and one alien/magnitude event alongside whatever alerts you already review. Write one policy proposal as a candidate, not as a silent constant edit. When you can freeze a pin, replay it.

The wiki gives the radar a model of the owner. The decision ledger records how that model was applied. The walk history records how the judgement was reached. Replay lets AI challenge the attention policy against the same past reality — without letting the future leak into the test, and without letting a single score hide a regression.

A quiet system becomes trustworthy when silence is a decision you can open, grade, sample, and re-run. Until then, low alert volume is not a result. It is an ambiguity.

References

  1. Hugging Face. “Security incident disclosure — July 2026.” — First-party account of autonomous-agent-driven intrusion into production infrastructure. Published 16 July 2026. https://huggingface.co/blog/security-incident-july-2026
  2. Russell Brandom / TechCrunch. “OpenAI says Hugging Face was breached by its pre-release models.” — Independent reporting used as the channel for OpenAI-attributed claims. OpenAI’s own disclosure page returned HTTP 403 in this writing environment and is not cited as a page read. 21 July 2026. https://techcrunch.com/2026/07/21/openai-says-hugging-face-was-breached-by-its-pre-release-models/
  3. Simon Willison. “OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened.” — Independent reconstruction of the public document stack. 22 July 2026. https://simonwillison.net/2026/Jul/22/openai-cyberattack/

Practitioner frameworks (author voice — not numbered inline)