Written and published before we know the result — so the interpretation can't bend to fit it.
Registered 2026-07-10: do plain-English mechanical entry rules produce better excess return over buy-and-hold than the same LLM committee trading freely, on the same market, timeframe, and cost model? Primary metric: median excess of mature rule arenas vs control excess, plus a majority test. Maturity thresholds are already met — the verdict will not be blocked on sample size unless something breaks in the final three weeks.
Would mean: under identical costs and market, mechanical gating extracted measurably more (or lost measurably less) than committee discretion. Evidence that discipline — not intelligence — is the operative variable at this timescale.
Would not mean: these rules "make money" (most arenas are negative in absolute terms), that they'd survive live execution, or that the result generalizes beyond BTC-perp 30m. One registered comparison, one answer.
Would mean: the committee's discretion added value over rigid gating — our prior was wrong, and we'll publish that sentence in exactly those words.
Would not mean: "AI can trade" — the control is also negative against buy-and-hold in our running data. Winning a relative contest between two losing strategies is a finding about mechanism, not a product claim.
Would mean: gating decisions through rules neither helped nor hurt versus free discretion — consistent with our audits showing directional IC ≈ 0 for both. The most likely reading: at this timescale, neither has edge to allocate.
Would not mean: "the experiment failed." A tie between two well-measured arms is a publishable answer. Most platforms simply never show you their ties.
Would mean: a registered maturity threshold was missed (e.g. arenas paused in the final weeks). The verdict publishes as exactly that — no extension, no loosened criteria, because prereg #2 registered no extension clause.
Would not mean: a hidden bad result. The per-arena numbers publish either way; only the registered aggregate claim goes unanswered.
Two material events occurred mid-experiment and are logged append-only: an exit-parameter change on 2026-07-23 (applied to both groups simultaneously — the comparison stays like-for-like) and a time-stop implementation deviation affecting both groups equally, fixed but held inactive until after the verdict so the experiment's behavior never changes mid-run. Both will appear in the verdict post's disclosure section.
Verdict due 2026-08-21 · experiment calendar · whatever it says, it ships.