aiarena· blog·experiments·2026-08-01

Pre-Registered: A Regime Filter We Refuse to Call Profitable (Yet)

Registered 2026-08-01 · verdict 2027-08-01 · criteria frozen before the first forward bar.

Read the label first. This experiment is registered as post-screen, risk-management-only. We are not claiming this strategy makes money, beats buy-and-hold, or has a verified Sharpe advantage. Our own forensics partially discredited its historical screen — details below, because that's the point of this platform.

The rule (all of it)

If yesterday's daily close > SMA200 → hold long. Otherwise → hold nothing. Executed at next day's open. No leverage. BTC and ETH USDT-perpetuals, paper-traded with real funding charged every period. That's the entire strategy. Two arenas: BTC · ETH.

What the historical screen showed — including what's wrong with it

SegmentWhat it said
Untouched perp 2019-2020 (real funding)Net +185%, drawdown −24% vs −60% for buy-and-hold, Sharpe 2.38 vs 1.85 for a same-exposure benchmark. Passed three frozen kill-tests (untouched history, 2× costs, parameter neighborhood 10/10 positive).
Jackknife auditDelete March 2020 and the risk-adjusted advantage vanishes (drawdown improvement 59.6% → 0.0%, Sharpe gap +0.53 → −0.36). The screen's edge is essentially one avoided crash.
2017-2020 spot reference (2018 bear, 2017 mania)Sharpe below buy-and-hold on both assets. Only the drawdown reduction (21–37%) survived. Spot, no funding — reported for regime coverage, not merged with the perp results.

Per our reviewer's wording, which we adopt verbatim:

"Historical screening showed drawdown reduction, but its risk-adjusted advantage was concentrated in March 2020 and did not persist in the longer 2017–2020 spot reference. This forward experiment tests whether any risk-management benefit generalizes; the historical screen is not treated as confirmation."

Registered criteria (frozen today, verbatim from protocol v1.2)

Verdict rules — 12 months, no extensions

Why register something this weak?

Because the alternative is what everyone else does: run it quietly, publish it only if the next 12 months happen to include a crash it dodges, and call that a track record. We'd rather freeze the criteria now, disclose that the screen leans on a single month, and let the forward data speak. The one thing that survived every segment, every parameter, and every audit was boring: it cuts drawdowns and stays positive after costs. Whether that generalizes is exactly what the next 12 months will answer.

Verify

Registered 2026-08-01 · first verdict eligible 2027-08-01 · how we read verdicts