EXP-000004
TYO SCORE: worst-month factor and drawdown cap
Hypothesis
Scaling consistency by the worst month, and capping the total by drawdown, restores discrimination without hand-tuning per system.
Controlled change
Consistency × max(0.5, 1 − |worst month|/100); total capped at 40/55/70 for DD ≥ 60/45/30%.
Dataset & conditions
All 14 systems, header + monthly series.
Before / After
Before
- Score range before the corrections
- 60–86, JAYRO at 60 on a 70.8% DD
After
- Score range after
- 40–84, JAYRO capped to 40
Validation
Out-of-sample, walk-forward, Monte Carlo and forward results are shown only where they exist as data. None exist for this entry.
Decision
RINA (2nd-highest profit factor, 56% DD) lands at 55 and JAYRO at 40 — the ordering now reflects what the score claims to measure. Caps are disclosed on every page they bind.
AI involvement
- Model
- Claude (Anthropic)
- What AI did
- Drafted the worst-month factor and the cap ladder; the thresholds (40/55/70 at 60/45/30% DD) were reviewed and approved by a human.
- Human review
- ✓ Reviewed and approved by a human
AI-assisted entries are published only after human review; entries generated by AI without that review are withheld by the build pipeline. AI does not predict markets, and no entry claims otherwise.
Notes
Model documented in docs/TYO_SCORE_MODEL.md; superseded by V2 in docs/TYO_SCORE_V2_MODEL.md.
Backtest improvement does not imply future improvement. The more experiments run, the more likely some succeed by chance — this log exists partly so that multiple-testing risk stays visible instead of hidden.