EXP-000004
TYO SCORE: worst-month factor and drawdown cap
Hipotesis
Scaling consistency by the worst month, and capping the total by drawdown, restores discrimination without hand-tuning per system.
Perubahan terkontrol
Consistency × max(0.5, 1 − |worst month|/100); total capped at 40/55/70 for DD ≥ 60/45/30%.
Dataset & kondisi
All 14 systems, header + monthly series.
Sebelum / Sesudah
Sebelum
- Score range before the corrections
- 60–86, JAYRO at 60 on a 70.8% DD
Sesudah
- Score range after
- 40–84, JAYRO capped to 40
Validasi
Hasil OOS, walk-forward, Monte Carlo dan forward hanya ditampilkan bila datanya ada. Entri ini tidak memilikinya.
Keputusan
RINA (2nd-highest profit factor, 56% DD) lands at 55 and JAYRO at 40 — the ordering now reflects what the score claims to measure. Caps are disclosed on every page they bind.
Keterlibatan AI
- Model
- Claude (Anthropic)
- Yang dikerjakan AI
- Drafted the worst-month factor and the cap ladder; the thresholds (40/55/70 at 60/45/30% DD) were reviewed and approved by a human.
- Tinjauan manusia
- ✓ Ditinjau dan disetujui oleh manusia
Entri berbantuan AI hanya dipublikasikan setelah tinjauan manusia; entri buatan AI tanpa tinjauan itu ditahan oleh pipeline build. AI tidak memprediksi pasar, dan tidak ada entri yang mengklaim demikian.
Catatan
Model documented in docs/TYO_SCORE_MODEL.md; superseded by V2 in docs/TYO_SCORE_V2_MODEL.md.
Perbaikan backtest tidak berarti perbaikan masa depan. Makin banyak eksperimen, makin besar peluang sebagian berhasil kebetulan — log ini ada sebagian agar risiko multiple-testing tetap terlihat.