tkalthe machine room

The Machine Room

tkal is an analyst operation I run with a stack of research tools I've built on top of Claude. This is the room those tools run in — a live build log of what I've built as Claude got more capable, and a weekly operator's log that grades the work in public. The premise is the scorecard's: wins here, losses here, louder.

It runs on one bet: that a set of narrow, disciplined tools — each built to keep receipts and grade its own calls — beats a single clever chatbot that talks a big game. I build and direct them; Claude does the analytical heavy lifting inside the guardrails I set. The build log below charts how the operation grew, and the pattern is deliberate: each jump in what Claude could reliably do unlocked the next piece. That's clearest on the trading side. Getting from a semiconductor read I ran by hand to an engine that can screen, size, place, and manage a real trade on its own took genuine advances in Claude's reliability and tool use — the difference between a model that only suggests and one you can trust to act inside strict limits. It began in early 2026 as hand-run "layer prompts" for semiconductors, and grew, model release by model release, into the research-and-trading operation below.

live   the build log · what got made as the models advanced

From the first hand-run semiconductor prompts in early 2026 to today's agentic trader. Each build is tagged with the AI capability that made it possible. Updated as new pieces ship.

Sep 22, 2026 · this week
The first option the engine bought and sold inside one session closed +$610 — and the hard part was getting a well-formed order out of the building
A morning in which every refusal was filed as “no opportunity.” By 08:00 PT both sleeves had screened five times and taken nothing, and the engine's own digest called all of it an absence of opportunity. None of it was a market read. On the stock side, the check that asks whether a name has breaking bad news was pointed at a browser tool the scheduled runtime does not have, so the answer was always unmeasured and an unmeasured answer blocks — correctly. On the options side, one component spells a direction BULLISH and the next one compares it against the literal call, so five candidates died at a gate whose name said the market had given no direction. Three more options runs, three more doors: a payload field the gate required and the broker's own tool forbids, a trading window two files disagreed about, and a per-order premium cap carrying the sleeve budget in the wrong field. Each fix moved a proposal one step further down the pipe, and the next fault was whatever it hit.
Then it traded. At 08:49 PT the options tick passed LRCX on conviction 3 of 3 — two screeners converging on it as a ranked bullish pick, flow with nothing contradicting it, and its sector rotating in — and bought two $305 calls expiring in three days at $6.05, with the stock at $302.99. Delta 0.47, spread 4.2% against a 12% ceiling, open interest 417 against a floor of 250, $1,210 of a $1,500 sleeve. It filled at 08:55. At 10:30 it marked $7.825, a whisker under the trail-arm threshold, and held. At 12:45 the take-profit rule fired on a mark of $9.225 and it sold at $9.10 with the stock at $309.86: +$610, +50.4%, three hours fifty-four minutes, and within 1.4% of the best mark the position ever printed. The stock moved 2.3%. The honest caveat is the timing — that name was sized and ready at 07:07 PT, so the repairs, not the signal, chose the entry price. Three cleverer exit rails were watching and disagreed with each other and with the blunt rule that actually traded; all three are still shadow-staged. The sleeve is now three winners in fourteen closed trades. The full build note →
unlocked by — a model that will stop on a contradiction instead of routing around it. The run log that hit the direction gate wrote down that the gate's name looked misleading and this was probably a wiring fault, and then it halted rather than inventing a direction to satisfy the check. Every one of the eight repairs started from a refusal the engine reported honestly about itself. The capability that mattered was not generating the fix; it was declining to fake the input.
Sep 15–20, 2026
Three systems that could not tell “nothing happened” from “I could not tell” — and the two keeping the records were the worst of them
The halt learns to end. The circuit breaker that stopped the engine on seventeen of the week's twenty-six runs took its leading reason from a reading computed once, in the morning — so a weak 6am print bought a whole session, and the release could not arrive before the next 6am. Asked on Sep 15 whether it should just run more often, the answer was to measure first: a shadow job took the reading every thirty minutes and wrote down what it would have said, changing nothing. Five sessions decided it — the raw reading flipped eight times, a version requiring three consecutive agreeing reads flipped twice, and on Sep 17 the confirmed version would have released at 07:57 PT while the daily file held to the close. It went live Sep 18 as a release leg only: it can lift a halt mid-session and stop a lifted one re-engaging on the same stale daily number, and it cannot create one. Intraday engage was deliberately not built — the ask was fewer halts. Rollback needs no code change. 198 governor checks, 24 of them new.
Two graders with mirror-image carry bugs. Sourcing the weekly grade turned up two figures that traced to a source file and were still wrong. The equity grader listed its windows rollup among the keys it copies forward untouched — and nothing anywhere produced it, so it had been photocopied since it was first written by hand, reporting a 17-trade lifetime while the block beside it in the same file read 20, with the payoff ratio and a −$195.18 drawdown computed off the stale slice against a ledger summing to −$196.35. The options grader had the exact inverse: it built its output from scratch and wrote it out without reading the file it replaced, erasing the weekly review's written analysis every night. That file had already diagnosed itself, in a sentence the next run was about to delete. Both fixed, both proven by running the old and new versions side by side, and both committed — the first time either grader has been under version control. The full build note →
unlocked by — nothing about the model, and that is the entry's own lesson. What mattered was answering “should this run more often?” with a job that measures instead of a change that assumes, then shipping only the half the measurement supported. The release leg runs no model at all. What a model did was read five sessions of shadow rows against the live governor, and — on Sunday — reconcile two blocks of the same JSON file against each other, which is where both grader bugs had been hiding in plain sight for weeks.
Sep 10–11, 2026 · this week
Three safety systems, from prose to code in two days — and the sniffer said "ok" for two hours while doing nothing
Three systems sit between the trading engine and a new position: the flow sniffer, which watches the options tape for unusual single-name buying and asks whether news already explains it; the EMA regime read, which sets the size dial from where the two big index funds sit against their moving averages; and the risk governor, the circuit breaker that writes one file saying whether the engines may open anything, and in which direction. On Thursday and Friday all three went from prompts a model re-read every morning to scheduled jobs that run no model at all. Each was audited first by a second model reading the live files, and each had the same defect: the rule lived in a sentence, its numbers lived in several disagreeing copies, and a missing input read as a passing one.
The sniffer was grading a score nobody had acted on — the run's log said 43.2 for an Adobe put, the ledger stored 34.2, and the gap was exactly one feature applied in one place and not the other. Its size haircut made a $50B+ company's ceiling 60 against an alert bar of 65, its $500,000 floor discarded 95.7% of a session's alerts, and its go-live test needed a comparison bucket that had n = 0 by construction. Version two: 45 modules, 195 checks, one immutable decision record with the score inside its id, a grading protocol hashed before any outcome, and a model that runs only when there is something to read. The governor engaged its long-only halt Thursday at 08:02 PT on live inputs, and its release test was satisfied at the same moment; the model producing it held the halt by a judgment it wrote down as unauthorised. Three readers carried three copies of "valid," and a stale regime file would have released a tier 3 halt the moment its evidence went missing. Now: one validator imported by both engines, the watchdog and the order hook at the broker's edge; a state machine that remembers each reason; every 30 minutes instead of four times a day. The regime read was arithmetically right to the cent and surrounded by four different sizes for the same tier — 0.0, 0.5, floored to 0.1, mapped to 1.0 — computed each morning by a task spending 27–41k output tokens to produce four averages. Now one contract, zero model calls, 0.4 seconds.
Friday, live. The sniffer's first two hours: every investigation refused by a missing manifest flag, zero model calls, status ok. Fixed 08:31 PT; then nine slots, 25 model calls, 23 cards — every half hour, from a system whose design says it almost always finds nothing. The governor's stricter input rule posted new risk blocked twenty minutes after the close, because the exchange had not yet published a VIX row for a fact every intraday run already held. Both false words came from a status that was asserted rather than derived, and both fixes make it derived. Not measured: whether the sniffer has any edge (2,806 observations, all unresolved), or whether the governor has ever saved a dollar. The full build note →
unlocked by — two models that can each hold a whole subsystem in view: one to read a live ledger row against a live log line and reproduce the bug from the difference, the other to carry a 45-module rewrite or a nine-stage repair through a single day without dropping the thread. The capability that mattered was not the writing. It was that the fix could be shaped as a contract file everyone imports — and that the model doing the importing could then be removed from the loop entirely, leaving it only where reading is the job.
Sep 2–5, 2026
The heaviest build week in this log — and the reason it was possible is the reason it nearly went wrong
Four days, four fronts: a ranking model that had quietly stopped responding to evidence, the last three safety gates on the trading engine, a written contract underneath the semiconductor pipeline, and the newsletter's figure work turned into a module. Then the part that belongs in a build log more than any of it — what it cost to direct four days of that pace.
Wednesday — the cyber ranking had frozen, and nothing could see it. The private-company model ranks about twenty security firms by likelihood of being acquired. A new check asked how much that list moves week to week: agreement between consecutive runs averaged 0.994 out of 1, four pairs at exactly 1.000, and one company held the identical probability — 0.761 — for eight consecutive weekly runs. Two defects underneath. The model's largest input was a ratchet, not a rate: across fourteen runs it moved up twelve times and down once while the sector's deal count halved (42 in February to 21 in July), with one category pinned at 0.95 for all fourteen. And five of thirteen run-pairs changed nothing at all — the block was copied forward rather than recomputed. Neither the accuracy score nor the calibration check could detect either, because both measure whether the numbers are right and neither measures whether they are new. Four versions shipped that day: a competing-risk term for companies that might list instead of sell, an AI exposure score conditioned on what kind of moat a company has rather than applied flat to the sector, the ranking metric that caught this, and the fix. A guard now refuses to write the run unless something moved, the numbers match the formula, and the series is not up-only over six runs. The first build of that fix was wrong in the most embarrassing direction — it read two summary fields as if they were months, inverting the correction so every category would have gone up — and it was caught by running it, not by reading it.
Thursday — three earnings deep dives in an afternoon (Broadcom, HPE, Western Digital), each published and each wired to its paid edition. Volume is not the point; the point is that the converter that makes it possible went in the day before, replacing a hand step that had been the actual bottleneck on this vertical since it started.
Friday — the trading engine went live. Six gates stand between the engine and real entries, each defined by written criteria a checking script confirms, on the rule that a gate is a computed fact, not an assertion, and that no script may flip its own gate. Three were already closed. Friday closed the two that cannot be closed by reasoning — a rehearsal placing real orders under per-drill approval, and a canary: one real share, placed by the engine, protected, watched. The test suite went from 490 to 529 passing between the morning readiness check and the arming, with 18 of its 30 files written or rewritten across these four days. The switch was armed by hand at 10:35 PT and the engine placed its first live order at 12:39 ET: one share of Ford, $14.5399, with a protective order placed alongside it. For about an hour and a half that stop was live at the broker and absent from the engine’s own book — not an unprotected position, an unverifiable one — and what noticed was not the trading run but the half-hourly monitor built the previous week to look for exactly this. Also Friday: all 127 scheduled jobs swept after one was found firing seven to nine times a month, and the fix for it was not better wording but a preflight script that prints proceed, skip or stop from three deterministic checks and halts the run before the model reads anything at all.
Saturday — contracts, and charts. The semiconductor pipeline's layers had no shared definition of what they pass each other; they agreed because they were written by the same hand in the same fortnight, which is coincidence with a short shelf life, and it broke twice this month. They now share one schema, one date-aware resolver, a credit calculation computed by a single tool that fails the run if a prompt's copy disagrees with it, and a check that must exit clean before anything publishes a number. On the newsletter side the figure work became a module — flow, waterfall, stacked and trend charts, plus a generator that builds an income-statement diagram from a filing's own line items — and 31 existing deep dives were backfilled with one.
Now the part that is not a feature. This pace is only available because the assistant reads the surrounding system and volunteers what it finds there — and that same behaviour is what nearly lost the week. Every session ends with a short list of things it noticed on the way past: a guard that should be tightened, two jobs that disagree about where a file lives, a check that would have caught the thing just fixed. Every item is real and every item is small enough to say yes to. Take one and it opens with its own list. The suggestions are generated from what the model just read, not from what you sat down to do — locally true, globally undirected. Four conversations in, the work is three branches deep in repairs nobody chose and the original task is still open behind you. What makes it hard to catch is that each individual step was an improvement.
The receipt is the notes file. The assistant keeps standing notes on how this operation actually works — how a schedule behaves, which file is authoritative, what a flag really does. It was created Aug 31. Six days later it holds 66 entries, most of them corrections written after something turned out to behave differently from what the last change had assumed. That is not a record of things learned so much as a record of a loop running faster than the person supervising it can read. The rule that came out of it is about sequencing, not capability: deliver the finding first and let the operator choose the fix; never chain a second change onto the first. Every graded engine here already carries a version of that rule — a minimum sample before any parameter moves, tweaks routed through a weekly review rather than applied on sight. It had never been applied to the act of maintaining the system itself.
unlocked by — tool-use reliability long enough to hold a four-day build across four unrelated systems without losing the thread of any of them. That is the capability, and this is the first week it produced more work than I could review at the speed it arrived. The scarce resource stopped being the model's competence and became my attention deciding which of its correct observations to act on. The three things that actually held the week — the classifier that refused to arm the trading switch, the preflight gate that returns a verdict before reasoning starts, and the monitor that found the share the engine could not vouch for — are all the same object: a check that runs outside the reasoning, on the assumption that the reasoning will be persuasive and occasionally wrong.
Aug 31, 2026
The floor gave out — the entire fleet moved off the cloud and onto the Mac in a day
The worst outage this operation has had. Every scheduled job — all of it, every pipeline in this log — ran inside a hosted Linux machine in the cloud. On Sunday morning that machine stopped being able to start. Not slow, not erroring halfway: it could no longer create the account a job runs under, so no line of code ran at all. The error was the same every time, four attempts in the morning and five more in a second pass that afternoon, each one under a different randomly-named session:
useradd: cannot create directory /sessions/lucid-quirky-bardeen — exit status 12
Different names, identical failure, which is what turned it from a bad session into a standing bug on the provider's side — nothing fixable from here. The jobs themselves were fine. They could still read and reason; they simply had no floor to stand on. The clearest picture of what that costs is in the day's safety report, which was supposed to check whether the trading engine was clear to go live and instead came back with no green and no red — no data at all. That is the worst answer a safety check can give, because it looks like nothing happened.
So the fleet moved home. Claude Code runs on the Mac itself, with the same skills and the same prompt files the cloud jobs were reading, which is why a rebuild that should have taken a week took a day. What changed, specifically:
  • 106 jobs rehomed the same day, each re-created locally as <taskId>-local, in production mode, on the identical schedule. Two manual-only jobs were mirrored later that afternoon, and one watchdog created after the roster snapshot — 110 running locally by evening. Verified rather than assumed: every local schedule checked against the cloud registry, every file path a job refers to confirmed to resolve.
  • The local jobs are pointers, not copies. Each one is a thin wrapper that reads the same canonical prompt file the cloud version read. This is the rule the last fortnight of audits was written in blood over: two copies of one job are two jobs that will quietly disagree, and the disagreement is invisible until it costs something.
  • Twelve jobs were deliberately left behind — the entire family that can place a broker order. Not caution for its own sake: run from this folder, the order guard denies every order, and that includes the protective exits that close a losing position. A trader that can neither open nor close is worse than one that is switched off, because it looks alive. They stay dark until they can run under a separate, gated launcher whose ten promotion checks are all still false.
  • The order guard became a hard veto. It now sits in front of every broker call any session in this folder can make and blocks it before it is sent, rather than relying on each job to check a flag it might not read.
  • Move, not copy. The cloud versions have to be switched off before that machine recovers, or the day it comes back every job fires twice and every post is published twice.
  • The plumbing got fixed on the way past: the Python interpreter pinned to one version rather than whatever the shell found; the machine's security certificates repaired system-wide; the chat poster taught to split anything over the platform's length limit instead of dropping it silently; and test runs now route to the test channel automatically rather than by remembering to.
The honest cost: local jobs only fire while the app is open on the Mac. That is a real downgrade from a machine that never sleeps, and it is the next thing to fix. The trade was made anyway, because a fleet that runs when the laptop is open beats a fleet that does not run at all — and because the failure exposed the sharper problem, which is that 110 jobs sat on one machine nobody here controls.
unlocked by — Claude Code being the same thing locally that it is in the cloud. The port took a day rather than a rebuild because every job in this log had already been reduced to a prompt file pointing at other files — the scheduler was a trigger, not the system. Portability was never a design goal; it was a side effect of the audits, and it is the reason a total infrastructure failure cost a Sunday instead of the operation.
Aug 27–29, 2026
No new pipeline — a four-pass loop that reads the system instead of adding to it
Nothing new shipped this week. What changed is how the prompts get written. They used to be written straight into the job that runs them, which is why so many of them quietly disagreed with each other. The newer ones come out of a loop that runs across two models, deliberately, because the failure being hunted is invisible from inside any single file:
  1. ChatGPT · first passAudit the system as it actually stands. Read the live jobs, the flags, the schedules and the files on disk, and report every place two of them state something that cannot both be true. No fixes, no rewrites — findings only, so the pass has nothing to defend.
  2. Claude · second passRebuild the skeleton. Take those findings and produce the structural map — what each piece is for, what it owns, what it reads and writes, which invariants have to survive the main file being unreadable. This is the pass that turns a list of contradictions into one shape, and it is where the actual repairs get written.
  3. ChatGPT · third passRe-audit against the skeleton. Same reader, new ground truth — but reporting only the summary of major changes, not another full inventory. A second exhaustive list would bury the diff; the point of the third pass is to confirm the shape held and name what moved.
  4. Claude · fourth passAudit ChatGPT's first evidence run, then merge it in. The daily semiconductor evidence sweep moved to ChatGPT last week, so its first output had to be checked before anything downstream could be built on it. The quality layer is written adversarially — it does not trust the collector, it verifies it against the sources — and it archives the raw file before touching a line, so the pre-audit version survives whatever the audit does. First graded run, Aug 28: PASS — every item retained, zero pruned, zero backfilled, zero duplicate URLs, and cross-run dedup confirming a YMTC capacity target was genuinely new while the earlier Nvidia and SK hynix items stayed suppressed as already promoted. Only then does the validated file merge back over the original. Two honest caveats: that run carried a single evidence item, so a clean pass proves the pipe works and nothing about the collector yet; and the weekly job that rolls the daily files into one has a valid schedule and a null last-run date — it has never fired once since it was created.
Neither model grades its own work, which is the whole reason for the split. The major changes it produced, in summary:
  • The safety net was unplugged. Long options have no protective order at the broker — the stop is software — and the monitor that checks it was reading the wrong folder, so it never once saw what had been paid for a position. Fixed Thursday; 10:03 ET Friday it closed the Intel put on a line it could not read 24 hours earlier.
  • A halt that stopped the brakes. A flag named for turning off chat posts disarmed every exit, trail and synthetic stop on live positions. The rule now, written identically in five places: a halt blocks new risk, never a risk-reducing exit.
  • The options sleeve got a master entry switch — its first ever, defaulted off and failing closed on anything that is not an explicit true. Both sleeves are now frozen for new positions and fully open for closing them.
  • One canonical folder, one calendar, one reader. Market holidays derived from the exchange's rules rather than typed into a list; a position reader that understands all three file schemas and reports could not read this as distinct from nothing is held; every scheduled job reduced to a pointer.
  • Three of six entry gates closed, with pass criteria written down for the first time and computed by a checker built so it cannot flip its own flags. One of them was a setting marked true since the file was written and read by no code anywhere.
  • The grader stopped trusting itself. It had been computing losses from the price an order was sent at rather than filled at, and counting its shadow exit rules without recording them — so no rule could earn promotion no matter how right it was. Broker fill is now authoritative; the counterfactuals get a scorecard.
unlocked by — two models good enough to point at the same system from opposite ends. One reads for contradictions without the ability to rationalise them away, because it did not write the thing; the other holds an entire tree in view and notices what exists only in the space between files — two copies of one job disagreeing about where state lives, a flag set to true and read by nothing, a guard whose default is inverted from its sibling's, a switch whose name and whose scope had drifted apart. None of it is visible in any single file, none of it throws an error, and every one was found by reading rather than by anything failing.
Aug 20–23, 2026
The fleet splits across two models — and gets an auditor to police the seam
The two heaviest daily gathering jobs — the semiconductor evidence sweep and the demand-evidence sweep — moved off the scheduler and onto ChatGPT, on the logic that wide, shallow collection shouldn't compete for the same budget as the layers that actually have to reason. Their outputs are now the freshest artifacts in the cascade (newest Aug 22, one day old, 53 and 51 files). Splitting a pipeline across two runners creates two silent failures, though: a step left switched on in both places runs twice and overwrites its own evidence, and a step nobody picks up stops producing while every gate downstream keeps reporting waiting on upstream. So the seam got a contract — a declared file naming, per step, who runs it now (chatgpt/cowork/manual/none), what it writes and reads, and how stale its newest output may be — plus a read-only auditor that checks the contract against the files on disk and the live scheduler. First run, Aug 23: the handoff is clean, and it surfaced 18 failures nobody was looking at, including one daily step declared live and silent for 58 days.
unlocked by — a fleet large enough that no one can hold its wiring in their head — and models cheap and reliable enough that the right answer is to split the work across two of them and write down which one owns what.
Aug 12, 2026
The before→after loop closes — a pre-print call, graded after the print
The anticipation engine ran its first full loop: freeze a dated call before a company reports, then grade it blind after. The debut case was Nebius — the day before its print it logged a read (direction a coin flip, the ~20% move the options market was pricing judged too rich to pay for → stand down), and after a contained +16% report it scored itself: direction correct, trade a wash, graded on separate scorecards. In the same run it planned the operation's first real defined-risk debit spread — buy one call, sell a higher one, loss capped by design — on Applied Materials into the Aug 13 close. It's the exact structure the live options sleeve keeps losing without; the job now is to move it from the lab into the room where real orders get placed.
unlocked by — agents reliable enough to hold a claim frozen across a real-world event and re-grade it blind — separating "was the read right" from "would the trade have paid" without a human scoring it by hand.
Aug 11, 2026
The anticipation engine — betting before the print, and fading the obvious
A separate, quarantined engine that takes a small, defined-risk position into an earnings event to test whether the mechanics can call the reaction — the mirror image of the reactive trader, which only ever follows a move. Its signature rule came out of a live CoreWeave read: when a beaten-down name in a washed-out sector refuses to break down the way every gauge says it should, that refusal — backed by quiet accumulation — is a contrarian buy, not a warning (an oversold spring, or the sector being bought back with the name out front). It runs on zero real capital, freezes a dated, falsifiable call before each event, and grades the direction apart from the trade — a correct read can still lose to the post-print volatility crush.
unlocked by — agents reliable enough to run a full evidence→hypothesis→blind-grade loop unattended — and to encode a trader's non-linear intuition as a rule that has to prove itself before it risks a dollar.
Aug 9, 2026
The Machine Room launches
This section — a weekly, self-graded operator's log covering the newsletter's three verticals and the agentic trading engine. The system now reports on itself in public.
unlocked by — models reliable enough at multi-source synthesis to read a week of its own logs and write an honest, cited summary without hand-holding.
Jul 27–31, 2026
"The week in 3 minutes" — companion video
The weekly issue gets a companion video: the newsletter's read, narrated and cut into a short, embedded on-site and cross-posted to Substack. The written issue became a watchable format.
unlocked by — text-to-narration and automated video assembly maturing enough to turn a written issue into a clean clip on a schedule.
Jul 21, 2026
The auto-trader starts trading unattended
The Robinhood engine adds autonomous midday entries — a "tick" that re-screens, sizes, and manages a live cash sleeve across the whole session, then grades itself every night and feeds the lessons back in.
unlocked by — steadier tool use and function-calling — the reliability to let an agent act on real orders inside guardrails, not just suggest them.
Jul 2026
The quant engine — the math moves to Python
All the arithmetic — position sizing, stop levels, option pricing (Black-Scholes) and dealer-gamma/GEX maps — gets pulled into deterministic Python scripts (tick_compute.py, options_compute.py, gex.py). The AI decides what to do; the code does the math. "Keep the model out of the arithmetic" becomes a house rule.
unlocked by — the hard-won lesson that models should orchestrate and call tools, not do numbers in their head — reliable code execution beats an LLM guessing a figure.
Jul 2026
The earnings deep-dive vertical opens
tkal adds a dedicated earnings section — 30+ post-print write-ups reading each company's quarter against the demand/credit cascade. The coverage goes from weekly synthesis to name-by-name depth.
unlocked by — cheaper, faster inference making it viable to run many research write-ups on a schedule without the cost breaking the model.
Jun 2026
The auto-trader goes live (paper→real sleeve)
First real, self-graded trades on a small cash sleeve, with a nightly post-mortem writing to a cumulative-learnings file the next run reads. The self-improving loop gets its first live test.
unlocked by — larger context windows — enough memory to carry a full trade ledger and its own learnings from one run to the next.
Jun 2026
Real-time options X-ray (Unusual Whales)
The engine plugs into live market plumbing through Unusual Whales — options order flow, dark-pool prints, dealer positioning (greeks/GEX), max-pain and the market "tide." The screeners stop guessing at sentiment and start reading where the real money is actually positioned. (A second feed, public.com, and a fail-soft fallback were wired in alongside as cross-checks.)
unlocked by — reliable tool/connector calling — letting an agent pull structured, live data feeds on demand and reason over them, instead of scraping pages.
May 2026
tkal launches
Issue No.001 ships (May 27), built on a twelve-layer evidence cascade across semiconductors, cyber, and construction. The founding idea — dated, falsifiable, self-graded calls — goes live.
unlocked by — models finally good enough at structured, multi-step reasoning to run a full evidence-to-call cascade and grade it honestly.
May 2026
Reading the finfluencers (YouTube)
A new data source comes online: the engine pulls YouTube transcripts from options "furus" (finance influencers), extracts the specific trades they call, and tracks how those calls actually play out — turning hours of video into a structured, gradeable signal, and a contrarian check on the crowd.
unlocked by — cheap long-context transcription and summarization — long enough and cheap enough to distill hours of video into a handful of structured signals a day.
Apr 2026
The pipeline learns to grade itself (the L13 calibration layer)
A calibration-and-audit layer gets bolted onto the semiconductor cascade: every dated claim is logged to a predictions file, then resolved blind against what actually happened and scored for hit-rate and calibration. This is the moment the whole system's DNA shows up — act, grade, keep the receipt — the exact loop the trader and the newsletter run today.
unlocked by — models steady enough to pull falsifiable claims out of their own output and re-check them later, instead of a human writing the scoring by hand.
Mar 2026
The cascade deepens — credit, demand, and a first trade signal
The semiconductor stack grows past raw supply-and-demand into hyperscaler credit and end-customer layers (the L8HC credit registry, the L1D demand layer), and spins off its first tradeable output — an options-watch feed dated March 23. The read stops being an essay and starts pointing at a position.
unlocked by — longer, more reliable reasoning chains — enough to hold a dozen interlocking layers in one pass and keep them consistent.
Jan–Feb 2026
The origin — hand-run "layer prompts" for semis
Where it all started: a stack of hand-written prompts for the memory/semiconductor market, run one after another in Claude by hand — L1 gathers the evidence, L2 maps the states, L3 forms the beliefs, and on down the chain. No automation, no trading, no self-grading yet — just one person running a reasoning cascade to understand the memory cycle.
unlocked by — a model finally able to hold a long, structured analytical prompt and return disciplined, layer-by-layer output instead of one undifferentiated blob.
the weekly operator's log

Every week: the headlines that moved the verticals, the fleet graded, and the lessons it wrote down.

No.008 · a week with no trades, and the six runs that said the wrong thing about it $0.00 for sep 14–18: five sessions, twenty-six runs, zero fills on either sleeve. seventeen of those runs sat under an engine-wide halt whose leading reason was a single daily reading — one that could be set in the morning and, as written, only lifted by the next morning. the intraday release leg that ends that shipped friday and gets its first real week on monday. the thirteen trades the engine declined were graded anyway: net +1.42R left on the table, all of it from wednesday’s +2.99R, and three of the week’s four winners were blocked by a volume reading the gate did not actually have. underneath it the payoff ratio is still 0.56 against a design of 3, with one trade in twenty ever exiting on a profit rule. also: six runs filed “no opportunity” while all thirteen input feeds read stale and zero candidates were fresh — a starved screen wearing a quiet tape’s label. and the newsletter published a second consecutive headline call that never reached its own scorecard. Sep 20, 2026
No.007 · one minute early, every day — the clock that cost it the week −$4.86 for sep 8–11 across three closed trades, none of them winners. the biggest loser was a first solar scalp that cost $2.14 because it was carried overnight, and it was carried overnight because the day's last run fired at 15:47 ET against a 15:48 cutoff — permanently ineligible to close anything. that schedule was moved the next morning; the options sleeve's last run still fires at 15:45 and its end-of-day pass is the only thing that can close an option. also: two ford lots finally exited after five straight runs that decided correctly and placed nothing, the first exit this engine has ever routed end-to-end; an intel entry approved at 13:35 ET that the broker never saw; and 94% of the account's settled cash still reserved for the sleeve that has lost $2,168.20. plus the newsletter's own version of the same fault — a call published sep 7 with a sep 10 resolver that was never written into the scorecard. Sep 13, 2026
No.006 · it placed its first live order — and then spent 90 minutes unable to prove it was protected +$12.00 for aug 31 – sep 4, on the week's only closed trade, and a human placed it. the stock sleeve's entry switch was armed by hand at 10:35 PT friday after the last two gates closed that morning; three hours later the engine bought one share of ford on its own, and for about an hour and a half that share's stop sat live at the broker and unrecorded in the engine's own book — an unverifiable position rather than an unprotected one, caught by the half-hourly monitor and not by the run that placed it. the harder number is zero: four of the week's five sessions wrote no decision record at all, for a different reason each day, so every gate statistic in the playbook is now five sessions stale. plus the grader withdrawing its own seven-week-old argument for trading more, once it checked what the number was denominated in. Sep 6, 2026
No.005 · a halt that turns off the exits is not a safety control — and the payoff, not the hit rate −$505.59 across both sleeves for aug 24–28. three days reading the machinery found two things: the between-tick stop monitor had been hydrating from the wrong folder, so the premium stop was never evaluated on a single option — fixed thursday, and friday morning it closed the intel put on a line it could not see 24 hours earlier — and a flag named for chat posts disarmed every exit, trail and synthetic stop on live positions, which on a long option is the only protection there is. nothing was lost to either. both sleeves are now frozen for new entries, options for the first time. plus the arithmetic the stock ledger has been holding back: average win +0.73R against average loss −1.43R, a realized payoff of 0.51 against a design of 3.0 — at which the sleeve needs to be right 66.3% of the time and is running 31.2%. Aug 30, 2026
No.004 · the first two wins closed by rule — then it broke the rule that made them options booked +$590 on aug 18, its first ever winners (chevron +$415, exxon +$175), both closed on the +50% take-profit rail rather than a decision — sleeve now 2W–8L, −$1,683. four days later it bought a salesforce put with earnings six days inside its life: the crush guard's volatility field came back empty and the check simply didn't run. the breaker fails closed; that guard fails open. stocks flat a fifth week, every block structured this time. plus the week's system changes — two daily jobs handed to chatgpt, and the read-only auditor built to prove they're still running. Aug 23, 2026
No.003 · a stale clock froze the stock engine — and a defined-risk spread still got the side wrong stocks zero trades in five days, blocked by an expired timestamp rather than a real halt; options −$1,021 realised (msft −$602, nvda −$419) taking the sleeve to 0W–8L, then bought two calls that finally avoid the old mistake (no earnings inside them, a month of room, smaller tickets) — +$112.50 on the mark. the event engine's first real debit spread called applied materials bullish into a beat-and-raise and the stock fell 5%: capping the loss isn't the same as being right. Aug 16, 2026
No.002 · the live sleeve ignored its own lesson — but the new pre-print engine got it right stocks flat (no-play days); options −$1,024 on two more naked calls (msft −$605, nvda −$419). meanwhile the new binary-event engine froze a nebius read before the print and graded it correct, and planned its first defined-risk spread — the fix, one module over. Aug 12, 2026
No.001 · the auto-trader graded — a green stock week, a red options week stocks +$12; options −$555 on two calls (a GOOGL call that broke its stop, an NVDA call pulled before earnings) — what it learned, the plan, plus the three verticals. Aug 9, 2026