aiarena· blog·experiments·2026-08-01

How to Read the August 21 Verdict

Written and published before we know the result — so the interpretation can't bend to fit it.

TL;DR — On 2026-08-21 a frozen script will resolve our first pre-registered experiment: 12 mechanical-rule arenas vs 1 free-running LLM committee arena. There are exactly four possible outcomes. This post locks in, today, what each one would mean and — more importantly — what it would not mean. On verdict day we will link back here and add nothing but the numbers.

The question being resolved

Registered 2026-07-10: do plain-English mechanical entry rules produce better excess return over buy-and-hold than the same LLM committee trading freely, on the same market, timeframe, and cost model? Primary metric: median excess of mature rule arenas vs control excess, plus a majority test. Maturity thresholds are already met — the verdict will not be blocked on sample size unless something breaks in the final three weeks.

Our prior, on the record

We expect most price-only mechanical rules to lose to buy-and-hold in absolute terms — our own audits show the committee's directional skill is indistinguishable from zero, and rules share the same market. The registered question is relative: whether rule-gated discipline beats unconstrained committee discretion. On that, our honest prior is roughly a coin flip with a slight lean toward the rules — discipline usually beats discretion when neither has edge. Writing this down now means we can't claim we "knew it all along," whichever way it lands.

The four possible outcomes

1 · rules-win

Would mean: under identical costs and market, mechanical gating extracted measurably more (or lost measurably less) than committee discretion. Evidence that discipline — not intelligence — is the operative variable at this timescale.

Would not mean: these rules "make money" (most arenas are negative in absolute terms), that they'd survive live execution, or that the result generalizes beyond BTC-perp 30m. One registered comparison, one answer.

2 · control-wins

Would mean: the committee's discretion added value over rigid gating — our prior was wrong, and we'll publish that sentence in exactly those words.

Would not mean: "AI can trade" — the control is also negative against buy-and-hold in our running data. Winning a relative contest between two losing strategies is a finding about mechanism, not a product claim.

3 · no-significant-difference

Would mean: gating decisions through rules neither helped nor hurt versus free discretion — consistent with our audits showing directional IC ≈ 0 for both. The most likely reading: at this timescale, neither has edge to allocate.

Would not mean: "the experiment failed." A tie between two well-measured arms is a publishable answer. Most platforms simply never show you their ties.

4 · insufficient-sample

Would mean: a registered maturity threshold was missed (e.g. arenas paused in the final weeks). The verdict publishes as exactly that — no extension, no loosened criteria, because prereg #2 registered no extension clause.

Would not mean: a hidden bad result. The per-arena numbers publish either way; only the registered aggregate claim goes unanswered.

What we've already disclosed

Two material events occurred mid-experiment and are logged append-only: an exit-parameter change on 2026-07-23 (applied to both groups simultaneously — the comparison stays like-for-like) and a time-stop implementation deviation affecting both groups equally, fixed but held inactive until after the verdict so the experiment's behavior never changes mid-run. Both will appear in the verdict post's disclosure section.

How to check us

Verdict due 2026-08-21 · experiment calendar · whatever it says, it ships.