Insights · Research

Backtesting and validation

How trading strategies are tested on historical data, the biases that make backtests misleading, and the validation practices that separate real effects from noise.

What a backtest is

A backtest runs a trading strategy over historical data as if it had been trading at the time. It produces a simulated record of trades, positions and returns. Used well, it is the cheapest way to reject bad ideas. Used badly, it is the easiest way to convince yourself that a bad idea is a good one.

Common biases

Look-ahead bias
Using information that would not have been available at the moment of the decision, such as a closing price to decide a trade made during the day, or data that was revised later.
Survivorship bias
Testing only on assets that still exist today. Instruments that were delisted or collapsed disappear from the data set, which flatters results. This is a particular issue in digital assets, where many tokens have ceased trading.
Overfitting
Tuning parameters until the strategy fits past data closely. The model learns noise, and live performance falls well short of the test.
Unrealistic fills
Assuming every order trades at the mid price, ignoring fees, spread, queue position, partial fills and the market impact of your own orders.
Ignoring latency
Assuming decisions and orders happen instantly, when in reality data and orders both take time to travel.

The multiple-testing problem

If you test enough variations of a strategy, some will look excellent purely by chance. Bailey, Borwein, López de Prado and Zhu showed in the Notices of the American Mathematical Society that the number of configurations tried is essential information. Without it, a high historical Sharpe ratio says little about future performance. The practical lesson is to keep a record of every variation tested and to demand stronger evidence as that number grows.

Good validation practice

  • Hold back an out-of-sample period that is never used while developing the strategy, and look at it once.
  • Use walk-forward testing: fit on one window, test on the next, then roll forward through the history.
  • Model costs conservatively, including fees, spread, slippage and a penalty for market impact.
  • Stress-test on specific difficult episodes, such as sudden crashes, exchange outages and low-liquidity periods.
  • Prefer simple strategies with few parameters and a clear economic reason to work.
  • Run the strategy in paper trading or shadow mode against live data before committing capital, then start small.

Metrics that matter

Sharpe ratio
Average excess return divided by the volatility of returns. A standard measure of return per unit of risk.
Maximum drawdown
The largest peak-to-trough fall in the value of the strategy. It shows the pain an investor would actually have lived through.
Turnover
How much is traded relative to capital. High turnover makes a strategy more sensitive to costs.
Capacity
How much capital the strategy can run before its own trading erodes the edge.
Hit rate and payoff
How often trades win, and how large wins are relative to losses. Either can be low if the other is high enough.
Key point

A backtest can prove that an idea is bad. It can never prove that an idea is good; it can only fail to find a reason to reject it.

Sources and further reading

  1. Pseudo-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out-of-Sample Performance · Bailey, Borwein, López de Prado and Zhu, Notices of the AMS
  2. Staff Report on Algorithmic Trading in U.S. Capital Markets · U.S. Securities and Exchange Commission

This article is for general information and education only. It is not investment advice, and it does not describe or solicit any product or service.