Nothing in this entry made money and nothing lost any. It is about the machinery that decides whether the engine is allowed to try — and about how much of that machinery turned out to be a sentence a model was re-interpreting every run.
Each system had the same shape of problem. The rule lived in a prompt. The numbers the rule depended on lived in several copies — the producer's instructions, a filter document, each engine, a health check — and the copies disagreed. And when an input went missing, every one of the three read "missing" as "fine": a stale regime file meant full size, a missing volatility read meant the volatility safeguard had passed, a failed investigation meant a quiet tape.
Each got the same shape of fix. One contract file, imported by everything that reads the state, so there is exactly one definition of valid. Consumers load the file themselves and fail to a fourth state, UNKNOWN, which blocks new risk and touches no exit — an object the model passes in is recorded and ignored. The producer becomes a scheduled job on the Mac that runs no model at all. And the model is called only where reading is genuinely the job, which in this stack is one place: the sniffer's investigation, where it goes and reads the filings and the press releases and says whether the flow is already explained.
The sniffer scores each unusual options print for how much it looks like someone positioning ahead of news, follows the strong ones on paper, and grades them a week later. The audit found that the score it graded was not the score it had acted on.
On Thursday's tape the run's own log scored an Adobe put 43.2; the ledger row for the same alert stored 34.2. The gap is exactly the multi-day-accumulation bonus (+15) times the large-company haircut (0.60): selection scored with history, the ledger scored without it. So every one of the 3,442 graded rows was graded on a number nobody had acted on. Underneath that, four more. A size veto that could not be passed: the raw score was capped at 100 and then multiplied by 0.60 for any company over $50B, so its ceiling was 60 against an alert threshold of 65 — and 44 of the 50 records in the snapshot were above the cutoff. A floor that threw out the population: one full session carried 4,368 alerts with a median premium of $58,211; only 190 (4.3%) cleared the $500,000 collection floor, which also excluded 1,022 alerts on sub-$10B names, exactly where the scorer's small-company points live. Duplicates: 3,078 distinct ids across those 3,442 rows, 364 extra rows from a dedupe that checked the ledger once per batch and never rechecked within it. A go-live test that could not be passed: readiness required the high-score bucket to beat the low-score bucket, and the low bucket had n = 0 forever, because nothing scoring under 40 ever became a proposal to grade — after 153 graded shadow trades the log still said null. And the grader applied today's quote to every matured row, so a row that matured on Tuesday and was graded on Friday got Friday's price.
Version two is 45 modules of plain Python, standard library only, 195 checks. A scheduled job fires at :00 and :30 and applies the exchange calendar itself (holidays, early closes, daylight time), and no model decides whether a scan is due. A scan collects, dedupes, scores, and stops; a tick that starts no model records exactly zero tokens. The model runs only when the queue holds a fresh, changed, eligible candidate, capped at three per slot, one per ticker, $2 per slot, and its researcher cannot set a score, a threshold, a destination or an order — its numbers are checked against the evidence packet by code, and a fabricated one downgrades the verdict. The score now lives inside the decision's id: the same evidence with a different score is a second visible row you can diff, not a silent disagreement. The $500,000 floor is gone; the size veto is gone (identical evidence scores the same at $3T and at $1.5B, and the replay produced eligible candidates in all four size groups). The grading protocol was written down and hashed before any outcome was computed, with evaluation starting Friday so no decision graded under it came from a scorer that had seen its own results; a matched control group is drawn at detection time from the same session's population, bypassing the score veto entirely, so the comparison that could never run now can. A horizon mark, once written, cannot be rewritten. There is no broker path in any of the 45 modules. Found on the way: a quadratic bug that made replaying six sessions take 63.6 seconds and get worse daily. 3.5 seconds after. A sidecar job finishes grading the 333 old rows under the old rules, at each row's own horizon-date close.
The first scheduled fire was Friday 06:30 PT. For the first two hours every candidate the dispatcher tried failed with the word disabled, zero model calls, zero alerts. The installed manifest was missing one flag that authorises live spend, and the dispatcher counted the failures and still returned ok, exit 0, which the run summarised as ok. It was not authentication: a probe run under the scheduler got a clean answer for $0.019. Fixed at 08:31 PT: an investigator that cannot run now exits 1 with a health block naming why, and the rule in the operations file is that silence from the sniffer is evidence of a quiet tape only when its health says ok. After the fix, nine half-hour slots ran the model — 25 calls — and posted 23 cards across 22 tickers: 20 verdicts of watch, 3 of already explained by news; confidence low on 15, medium on 7, high on 1; scores 65 to 78; the three-per-slot cap hit in five of the nine slots. The design document says the sniffer "almost always finds nothing worth interrupting you about." Day one interrupted every half hour. Whether that is the tape or a provisional threshold is exactly what the frozen protocol exists to answer, and nothing moves until it has. Two honest gaps: the nine model slots logged as dispatcher-reported with no token count and no dollar figure, so the cost ledger holds no number for Friday's spend (the bounded live check on Thursday measured about $0.25 per investigation); and at the weekly checkpoint 2,806 observations sat under the protocol, every one unresolved, because the capture of horizon prices is not yet on a schedule.
The governor writes one file: halted or not, in which direction, and why. Both engines read it before opening anything. On Thursday morning, on live inputs, it engaged for the first time under a rule written in July — and the rule for letting go was satisfied at the same moment.
At 08:02 PT Thursday the structural-risk overlay came off its 61 floor to 65 (band COOL to NEUTRAL, four-week velocity 0 to +4, FLAT to DRIFTING_UP), which flipped the breaker's first condition true for the first time since the July 15 rule; with the regime at tier 2 and the S&P fund down three sessions and the Nasdaq fund down two, it engaged LONG_ONLY: no new bullish positions, bearish puts still eligible. Then the defect. Engage read band HOT or velocity rising; release read band back to NEUTRAL or COOL, or velocity improving or flat. A NEUTRAL band with rising velocity satisfies both. The model running the producer held the halt anyway and wrote in its own transcript that the specification did not cleanly authorise the judgment — the right outcome, with no rule behind it. And the overlay file's own caveat says a +4 velocity "sits inside the noise band of the scoring method itself" and that the overlay "should NOT size or gate real money on its own"; the producer's instructions read exactly that field to gate the breaker. Three more, reproduced at function level with nothing written: the options engine's Thursday dry-runs had no breaker object at all on three of four rows, because the model assembling the input forgot to attach it — UNKNOWN, policy warn, entries not blocked (dry-run, no order). Three readers carried three copies of the validity rule: both engines crashed on a timestamp with no timezone, so an otherwise-valid exit result was never emitted; an expired CLEAR whose time-to-live read "NaN" was CLEAR forever, because NaN compares false against everything; a halted field of empty-string was CLEAR; the mode word informational switched equity enforcement off entirely; and a stale regime file read as tier 1, which means a tier 3 halt would have released the moment its own evidence went missing. The order hook at the broker's edge never consulted the governor at all.
Stage 0 hashed every file to be touched and built a replay harness that runs real decision bundles through the engines, so each later stage is checked against what the engines actually saw and not a fixture. Then: one validator, standard library, exception-safe and deterministic, imported by both engines, the watchdog, and now the order hook at the broker's edge, which denies an entry under any halt and denies both directions on UNKNOWN. Both engines load the file themselves; an object the model passes is archived as ignored. Options on UNKNOWN now block, not warn. The producer is code: an explicit state machine that remembers each reason separately — missing evidence holds an active halt, cannot engage an inactive one, and never releases anything; the overlay leg releases only on the strict complement of its engage test and a newer file than the one that engaged, so an unchanged Monday file can never be re-read as "eased." It refuses to write a time-to-live shorter than its cadence plus an hour, keeps a durable outbox so a failed post is retried rather than lost, and posts on any change of halted, scope, or reasons with text that says which direction is actually blocked. It runs every 30 minutes with a 90-minute time-to-live, against four model runs a day at 180 before. Shadow ran one afternoon, 7 rows, 0 disagreements, and the three-day precondition was waived by me. A measurement script now exists that joins what each tick saw to what the blocked candidates did next, and refuses to report zero trades as benefit. Tests: 174 on the governor, 94 on the validator, 55 on the order hook. Since the cutover: 19 production writes, one scope change and one reason change, both posted with their message ids in the ledger.
Friday's fourth audit reproduced a case the Thursday rebuild had left open: an inactive governor with a missing volatility read graded PARTIAL and CLEAR, and the consumers accepted that as "the volatility safeguard was checked." It was not. Version 4.1 names four mandatory inputs: the regime, the index closes, the prior VIX close, the live VIX level. Any of them missing or stale makes the run INCOMPLETE, which every consumer reads as UNKNOWN, halted or not; the overlay is optional and only degrades the clearance. The accepted cost, written down: an intraday outage of the unofficial VIX quote blocks new risk for one 30-minute cadence. Then the rule's first afternoon. At 13:20 PT both index funds had closed up (S&P fund 764.29, Nasdaq fund 714.88), so the tape leg released and the streak reason came off; but the run graded INCOMPLETE and its post said new risk blocked, mandatory input missing, twenty minutes after the close, when nothing trades. Diagnosis: the governor calls Friday a completed session twenty minutes after the bell and then wants Friday's row from the exchange's official VIX history file, which the exchange publishes hours later. The prior close it needed was the same number every intraday run had already used; the observation was never missing, only the spreadsheet was late. Fix, same afternoon: a row exactly one session behind, read on that session's own date, is OK and flagged publication lag; the next pre-open run demands the official row, and two sessions behind is stale on any date. The 13:32 PT write graded COMPLETE. State at the close: halted, scope ALL, reason EMA tier 3 — and underneath that, the S&P fund closed Thursday 0.052% below its 50-day average and the Nasdaq fund 0.31% below. That breach blocks new positions in both directions, puts included. The audit called that too broad, and the response is a shadow policy, under which tier 3 alone would be long-only and a volatility spike would still be all, now written into the file and the ledger and applied by nothing, with both engines recording per candidate what it would have done. Zero differing rows so far, because no put has reached the breaker as its decisive gate. The governor's own benefit, measured on Sep 4–11: 27 deduplicated opportunities, 7 blocked by the market state, every row labelled insufficient. And a sensitivity read over 542 sessions: 17 tier 3 episodes, median engage depth 0.34% below the 50-day average, half released within two sessions, the index up over 12 of the 16 that closed. No buffer proposed. If one is, it gets tested forward, not fitted.
The regime read is the size dial: tier 1 is full size, tier 2 half, tier 3 defensive. The audit re-derived Friday's numbers by hand and they matched to the cent. Everything around the numbers was wrong.
Friday's state was tier 3: S&P fund close 757.83 against a 50-day average of 758.23, Nasdaq fund 708.69 against 710.89, both below, recomputed independently from the saved closes, all eight averages to the cent. Then the contract. The producer's instructions said tier 3 meant equity size 0.0; the filter document, revised Aug 30, said 0.5; the equity engine floored any multiplier it received at 0.1; the options engine mapped a zero to 1.0, full size. The options engine ran on the wrong regime on two of Thursday's three ticks — tier 1 at full size while the file said tier 2 at 0.7 — because the model assembling the input forgot to pass it; a missing, stale or malformed file read as tier 1 everywhere. The leadership test compared rounded values. And the producer itself was a model-supervised task that regenerated its own calculation script every morning: 28 to 54 assistant messages and 27,074 to 40,644 output tokens per run, to compute four moving averages. Beside it, an individual-stock bounce tracker with no producer and no consumer, a price cache in which 38 names held zero bars and 43 sessions were missing, and a July research file whose path test had leaked future volatility into its own stop and target.
One module now owns four things and nothing else: the arithmetic, the classifier, the validator, and the consumer policy. The tier rule is written once: either index at or below its 50-day average is tier 3; both above the 21- and 50-day with the Nasdaq fund leading on the unrounded stretch is tier 1; otherwise tier 2; the 8-day average is descriptive only. The multipliers are 1.0 / 0.5 / 0.5 for equity and 1.0 / 0.7 / 0.5 for options, and a multiplier of zero is not floored and not mapped: it is UNKNOWN, and UNKNOWN blocks new risk and touches no exit. Both engines load and validate the file themselves. At tier 3 an options entry needs a delta of at least 0.45 and two independent, agreeing classes of evidence from the candidate's provenance — the same data feed repackaged twice counts once, and conviction is not evidence. The producer is a scheduled job at 05:35 PT with an idempotent retry at 05:50: two REST calls, zero model calls, 0.4 seconds, cut over at 07:07:35 PT Friday with the old task disabled. Tests: 157 on the contract and producer, 71 on the engines' actual assembly paths, including the omitted-field and contradictory-field cases the audit reproduced. The bounce research was rerun with the leak fixed, on both the 73-name and the S&P 500 universes: no demonstrated edge at any horizon — the S&P five-day spread of +0.15% sits under the 0.65% the sample could even detect, so the predictor stays unwired, and "no demonstrated edge" is stated as what it is, not as proof of no edge. The cache was repaired: 72 of 73 names valid, one excluded by the corporate-action rule.
| system | producer before | producer after | model tokens per run | cadence |
|---|---|---|---|---|
| flow sniffer | model task, 10:30 + 12:30 PT | scheduled job, :00/:30, calendar-gated | ~440,000 reported (a measurement bug: bytes/4 of a 1.77MB ledger) → 0 on a scan; model only inside an investigation | 2 → up to 14 in-session scans |
| risk governor | model task, 4× a day | scheduled job, every 30 min | estimated → 0 | 4 → 17 fires, TTL 180 → 90 |
| EMA regime | model task, 05:35 PT | scheduled job, 05:35 + 05:50 retry | 27,074–40,644 output → 0 | unchanged |
Two words were wrong on Friday and they were wrong in the same way. The sniffer said ok for two hours because its dispatcher counted failures and returned a status it had not derived from them. The governor said blocked after the close because it demanded a spreadsheet the exchange had not published yet. Neither system had done anything dangerous; both had reported something false about themselves, and a false status word is the one output nothing downstream can check. The fix in both cases was the same: the word is computed from the thing it describes, a wrong one exits non-zero, and the receipt carries the reason.
Two more things belong in the record. The token budget found the duplicate the audit didn't. The premarket options scan returned skip on Thursday and again on Friday because its budget group ran at roughly twice its weekly cap — and the spender was not the scan. The old sniffer task and its local twin were charging 4.58M tokens a week between them against a 4.00M cap for the whole group, a cloud/local duplicate of exactly the shape the Aug 31 port warned about. The cap was not raised: the old task was already retired by Thursday's cutover, and the premarket scan came off the list of jobs the budget may skip, because it fires once a day and was 4% of its group. And the safety classifier refused several of these edits until I authorised each one — the order hook, running the governor from a session, the sniffer's prompt file. Same object as last week's note: a check that runs outside the reasoning, on the assumption the reasoning will be persuasive and occasionally wrong. It was, once: an edit that carried its own backup in the same refused call ran with no rollback point when the call was retried without it. The rule that came out of it is that the backup is its own step.
Not measured: whether the sniffer's detector has any edge: every weight is a provisional hypothesis and 2,806 observations are unresolved until horizon prices are captured on a schedule; paper execution has never produced a fillable row because this data key serves no per-contract option prices; the governor's benefit, which the measurement script labels insufficient on every row. Not proven: the EMA producer under its own scheduler, which fires for the first time Monday 05:35 PT — Friday's cutover was a hand-run of the same code. Not resolved: a halt on both directions for a 0.052% breach of a moving average, which the shadow policy will argue about in the ledger and not in production; and a fleet that is faster to guard than to act — equity scans at six fixed times a day, a governor every thirty minutes, a regime read once at dawn, so an early selloff can precede every signal that would have stopped the engine from buying it. Nothing changed cadence this week, and adding another gate would not fix timing.
Repair handoffs and cutover runbooks for all three systems, the four audit documents, the governor ledger (26 rows: 7 shadow, 19 production), the sniffer's tick log and checkpoint, the Discord delivery ledger (23 sniffer cards, 3 governor posts, 3 regime posts across the two days), and five test suites re-run for this entry: 195, 174, 94, 157 and 71 checks, all passing. Figures are as recorded in those files at the Friday close. Times are Pacific unless marked ET.