EXP-000004
TYO SCORE: worst-month factor and drawdown cap
仮説
Scaling consistency by the worst month, and capping the total by drawdown, restores discrimination without hand-tuning per system.
統制された変更
Consistency × max(0.5, 1 − |worst month|/100); total capped at 40/55/70 for DD ≥ 60/45/30%.
データセットと条件
All 14 systems, header + monthly series.
変更前 / 変更後
変更前
- Score range before the corrections
- 60–86, JAYRO at 60 on a 70.8% DD
変更後
- Score range after
- 40–84, JAYRO capped to 40
妥当性確認
アウトオブサンプル、ウォークフォワード、モンテカルロ、フォワードの結果は、データが存在する場合のみ表示します。本エントリには存在しません。
判断
RINA (2nd-highest profit factor, 56% DD) lands at 55 and JAYRO at 40 — the ordering now reflects what the score claims to measure. Caps are disclosed on every page they bind.
AI関与
- モデル
- Claude (Anthropic)
- AIが行ったこと
- Drafted the worst-month factor and the cap ladder; the thresholds (40/55/70 at 60/45/30% DD) were reviewed and approved by a human.
- 人間によるレビュー
- ✓ 人間がレビューし承認済み
AI支援の項目は人間のレビュー後にのみ公開されます。レビューを経ていないAI生成の項目はビルドパイプラインが公開を差し止めます。AIは市場を予測せず、そのような主張をする項目はありません。
備考
Model documented in docs/TYO_SCORE_MODEL.md; superseded by V2 in docs/TYO_SCORE_V2_MODEL.md.
バックテストの改善は将来の改善を意味しません。実験の数が増えるほど、偶然うまくいくものが混ざる確率は上がります——このログは、多重検定リスクを隠さず可視化するためにも存在します。