Honest writeups from building a multi-agent LLM trading arena — negative results included, on purpose.
21 public arenas, 131 trades closed this week, 2/16 mature arenas beating buy-and-hold. Auto-generated — losses included.
read →21 public arenas, 113 trades closed this week, 0/16 mature arenas beating buy-and-hold. Auto-generated — losses included.
read →21 public arenas, 186 trades closed this week, 2/16 mature arenas beating buy-and-hold. Auto-generated — losses included.
read →A daily SMA200 exposure filter on BTC/ETH perps, registered post-screen as a risk-management claim only. Our own jackknife showed the historical edge concentrates in one month — March 2020 — and we publish that up front. Criteria frozen, 12-month verdict, no early-confirmation path.
read →Our first pre-registered verdict resolves in three weeks. There are exactly four possible outcomes. This post locks in — today — what each one would and would not mean, plus our prior on the record, so the interpretation can't bend to fit the result. On verdict day we add nothing but the numbers.
read →Make a real strategy look good with one line of code and one choice of window — then freeze your own kill conditions and meet the untouched history that actually killed it. Built entirely from our archived research. No leaderboard, no deploy button; the only thing measured is whether you change your rules after seeing the result.
play →120 cells searched, 113 couldn't beat holding. The one candidate that could was made to run a gauntlet with failure conditions written before each test: untouched 2020-2023 history, 3.6× fee stress, a 9-cell parameter neighborhood. It survived — so it gets a pre-registered paper arena, a frozen hypothesis, and a 2026-10-11 judgment date. Registered prior: catastrophic drawdowns, and it should lose to buy-and-hold in bull markets.
read →One cell out of 54 made +44.7% over three years. Then we swapped spot for perp data, charged real funding, split the windows, and tried other assets — the edge died four times. The live arena keeps running to its pre-registered judgment date anyway, and every number ships in a downloadable evidence package.
read →Twelve mechanical-rule arenas vs one rule-free committee control. Instead of picking the flattering window after the fact, we declare the hypothesis, exact metrics, sample thresholds, and evaluation date now — plus a second experiment on whether nine independent LLM calls beat one batched one. Whatever comes out on 2026-08-21 gets published.
read →high.rolling(20).max() includes the current bar, so close > that is false on every bar that has ever existed. Zero out of a thousand. Nothing crashes. Live, your breakout never fires; in a backtest, the mirror image hands you tomorrow's data. Plus two more silent-false conditions, and the open-source linter that finds them.
Paper copy-trading and live exchange mirroring shared one position slot, so a live position could be opened and never closed. Twenty-seven orders failed in a row and nobody was told. Fixing the first bug would have double-counted every subscriber's return. And the scoreboard promising to show "when we lose to buy-and-hold" didn't show buy-and-hold at all.
read →A committee peaked +17% then gave back to +8% in a range-bound market. The fix wasn't code — it was a plain-English rule that trades fewer, higher-quality setups (a support/resistance pyramid). Plus 12 control arenas testing which entry rules actually have an edge.
read →The arena beat its own judgment-free baseline by 33 points — and the same dashboard shows a 55% win rate, one dominant trade, and an "immature" label. The honest teardown of a green curve, plus the first time the "smart" layer leads the live A/B.
read →Breeding overfits, equal weight beats "smart" weighting, IC and PnL can disagree in sign, and of ~94 factors exactly one survives a strict coin-and-time holdout (16 → 13 → 0). But every test ran on the bare scaffolding — the edge layer was never on the table.
read →Research and paper-trading. Not investment advice.