Baseline: measured 2026-09-26 on v0.18.0 (#826, R-based)
keel simulate --years 5 over the daily paper rule set (config.paperforward.yaml, 19 assets), 1.20% taker plus per-product slippage:
| pooled N |
win rate |
expectancy |
avg win / loss |
PF (R) |
positive assets |
| 220 |
25.0% |
+0.166 R |
+4.14 R / −1.16 R |
1.19 |
12 / 19 |
How far that is from break-even: the payoff ratio is 4.14 / 1.16 = 3.57, so break-even is a 21.9% win rate. The observed win rate is ~3 points above break-even. With a design effect of 2.58 (#427), n_eff ≈ 85, and at 80% power the detectable edge is ~13 points. The baseline is not distinguishable from zero, which is ADR 0006's descriptive regime.
Note: an 8-asset subset of this run (BTC, ETH, PAXG, SOL, XLM, LTC, ADA, LINK) reads +0.30 R over about 82 trades. That is a subset of the same run, not the baseline, and it must not be used as one.
Already in the rule, so not a candidate
The daily Turtle already gates on ADX(14) > 25, a choppy-regime filter and a higher-timeframe bias. "Regime filter = ADX > 25" is the baseline, not a treatment.
Hypothesis
Entries against a falling long-term trend are disproportionately false breakouts. Gating entries on the 200-day SMA raises pooled expectancy in R.
Arms, declared now and not tuned afterwards
- A.
close > SMA200 on the decision bar.
- B.
SMA200[t] > SMA200[t-5] (5-day slope positive).
- C. A AND B.
Everything else in the rule is unchanged.
Criteria (pre-registered)
- Same engine and data: the
simulate edge pass, same 19 products, same 5-year window, 1.20% fee plus per-product slippage, in R.
- Report every arm: pooled N, win rate, expectancy_r, PF(R), n_eff, and the detectable-edge sentence (ADR 0006). Per-product rows are diagnostics only.
- Retention floor: an arm below 100 pooled trades is reported but not evaluated. ADR 0006's floor, not a lower one.
- Multiple testing: three arms mean three trials. Each run records to the trials ledger with
--trial-decision diagnostic_only, and any "improvement" is stated against the deflated bar (keel research deflate).
- Out-of-sample: a filter that helps in-sample is walk-forward validated (
keel research walk-forward) before any promotion discussion.
- Honest null: "no arm improves on the baseline beyond its detectable edge" is a complete result.
The record goes in docs/experiments/, with the driver docstring carrying this pre-registration before the run.
Baseline: measured 2026-09-26 on v0.18.0 (#826, R-based)
keel simulate --years 5over the daily paper rule set (config.paperforward.yaml, 19 assets), 1.20% taker plus per-product slippage:How far that is from break-even: the payoff ratio is 4.14 / 1.16 = 3.57, so break-even is a 21.9% win rate. The observed win rate is ~3 points above break-even. With a design effect of 2.58 (#427), n_eff ≈ 85, and at 80% power the detectable edge is ~13 points. The baseline is not distinguishable from zero, which is ADR 0006's descriptive regime.
Note: an 8-asset subset of this run (BTC, ETH, PAXG, SOL, XLM, LTC, ADA, LINK) reads +0.30 R over about 82 trades. That is a subset of the same run, not the baseline, and it must not be used as one.
Already in the rule, so not a candidate
The daily Turtle already gates on ADX(14) > 25, a choppy-regime filter and a higher-timeframe bias. "Regime filter = ADX > 25" is the baseline, not a treatment.
Hypothesis
Entries against a falling long-term trend are disproportionately false breakouts. Gating entries on the 200-day SMA raises pooled expectancy in R.
Arms, declared now and not tuned afterwards
close > SMA200on the decision bar.SMA200[t] > SMA200[t-5](5-day slope positive).Everything else in the rule is unchanged.
Criteria (pre-registered)
simulateedge pass, same 19 products, same 5-year window, 1.20% fee plus per-product slippage, in R.--trial-decision diagnostic_only, and any "improvement" is stated against the deflated bar (keel research deflate).keel research walk-forward) before any promotion discussion.The record goes in
docs/experiments/, with the driver docstring carrying this pre-registration before the run.