Backtest Lab
Why we publish models that failed
A note on why the discarded versions of a system belong in its documentation, and what a reader should look for in any published backtest.
PLACEHOLDERSample content — this description is structural placeholder copy and will be replaced with the finalised product documentation.
Hypothesis
A published result is only interpretable when the reader also knows what was rejected on the way to it.
Method
We compare how a result reads with and without its rejected predecessors and its test conditions attached.
Result
Without conditions and rejected versions, a result is indistinguishable from a curve that was fitted to the data it is being shown on.
Conclusion
Publish the conditions, publish the failures, and let the reader judge. Anything less asks for trust that has not been earned.
A backtest is a measurement, and like every measurement it is meaningless without its conditions. The same strategy, tested on different tick data, with a different spread assumption and a different modelling method, will produce results that disagree with each other by more than most people expect.
What a reader should demand
- The exact test period, and whether it includes the period the idea came from.
- The data source and its modelling quality.
- Spread and commission assumptions, stated as numbers.
- How many parameter combinations were tried before this one was shown.
- What the version before this one did, and why it was discarded.
Why the failures matter most
A model that survived ten rejected predecessors is a different object from a model that worked the first time. The rejections are the evidence that the surviving version was selected for a reason rather than found by searching until something looked good.
This article is a methodology note. It contains no performance claim and no result for any TYO system.