What a backtest is
A backtest runs a trading strategy over historical data as if it had been trading at the time. It produces a simulated record of trades, positions and returns. Used well, it is the cheapest way to reject bad ideas. Used badly, it is the easiest way to convince yourself that a bad idea is a good one.
Common biases
- Look-ahead bias
- Using information that would not have been available at the moment of the decision, such as a closing price to decide a trade made during the day, or data that was revised later.
- Survivorship bias
- Testing only on assets that still exist today. Instruments that were delisted or collapsed disappear from the data set, which flatters results. This is a particular issue in digital assets, where many tokens have ceased trading.
- Overfitting
- Tuning parameters until the strategy fits past data closely. The model learns noise, and live performance falls well short of the test.
- Unrealistic fills
- Assuming every order trades at the mid price, ignoring fees, spread, queue position, partial fills and the market impact of your own orders.
- Ignoring latency
- Assuming decisions and orders happen instantly, when in reality data and orders both take time to travel.
The multiple-testing problem
If you test enough variations of a strategy, some will look excellent purely by chance. Bailey, Borwein, López de Prado and Zhu showed in the Notices of the American Mathematical Society that the number of configurations tried is essential information. Without it, a high historical Sharpe ratio says little about future performance. The practical lesson is to keep a record of every variation tested and to demand stronger evidence as that number grows.
Good validation practice
- Hold back an out-of-sample period that is never used while developing the strategy, and look at it once.
- Use walk-forward testing: fit on one window, test on the next, then roll forward through the history.
- Model costs conservatively, including fees, spread, slippage and a penalty for market impact.
- Stress-test on specific difficult episodes, such as sudden crashes, exchange outages and low-liquidity periods.
- Prefer simple strategies with few parameters and a clear economic reason to work.
- Run the strategy in paper trading or shadow mode against live data before committing capital, then start small.
Metrics that matter
- Sharpe ratio
- Average excess return divided by the volatility of returns. A standard measure of return per unit of risk.
- Maximum drawdown
- The largest peak-to-trough fall in the value of the strategy. It shows the pain an investor would actually have lived through.
- Turnover
- How much is traded relative to capital. High turnover makes a strategy more sensitive to costs.
- Capacity
- How much capital the strategy can run before its own trading erodes the edge.
- Hit rate and payoff
- How often trades win, and how large wins are relative to losses. Either can be low if the other is high enough.
A backtest can prove that an idea is bad. It can never prove that an idea is good; it can only fail to find a reason to reject it.
Sources and further reading
- Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance · Bailey, Borwein, López de Prado and Zhu, Notices of the AMS
- Staff Report on Algorithmic Trading in U.S. Capital Markets · U.S. Securities and Exchange Commission
This article is for general information and education only. It is not investment advice, and it does not describe or solicit any product or service.