Testing a strategy on historical data is the only way to know whether it has an edge. It is also an almost perfect machine for generating false confidence, because the past is fully known and the human mind is extremely good at finding patterns in it.
The five ways a backtest lies
| Bias | What it is | How to defend |
|---|---|---|
| Overfitting | Tuning parameters until the equity curve looks beautiful. With enough knobs you can fit any history and predict nothing. | Minimise parameters. Verify that neighbouring values also work — if only a 47-day average succeeds, you found noise. |
| Look-ahead bias | Using information that was not available at the time — buying at today’s open using today’s close, or a fundamental figure before it was published. | Rebuild the test so every decision uses only data that existed at that timestamp. Indian results are often filed weeks after quarter end. |
| Survivorship bias | Testing on today’s NIFTY 50 across ten years. Those are the companies that made it — the delisted and collapsed ones are missing. | Use point-in-time index constituents, or accept that your result is optimistic by a known and large margin. |
| Ignoring costs | Omitting brokerage, STT, stamp duty, GST, slippage and the bid-ask spread. | Include all of them. A high-frequency strategy that looks profitable gross is very often loss-making net. |
| Period selection | Testing only on 2020–2024, a period in which almost every long strategy worked. | Include at least one severe bear phase — 2008, 2011, 2018 smallcaps, March 2020. |
Walk-forward testing
The single most valuable defence. Rather than optimising over all your data and admiring the result, split it.
- 1Develop on the first portion
Use, say, 2010–2017 to design rules and choose parameters. This is your in-sample data and you are allowed to look at it as much as you like.
- 2Freeze the rules completely
Write them down. No adjustments after this point — that is the entire discipline, and it is harder than it sounds.
- 3Run once on data you have never touched
2018–2025 is your out-of-sample test. You get one attempt. If you go back and tweak after seeing it, that data is now in-sample and worthless as a test.
- 4Expect degradation, and judge by how much
Out-of-sample results are almost always worse. A modest fall is normal. A collapse from excellent to breakeven means you fitted noise.
Look past the total return
A headline CAGR tells you very little about whether you could actually have traded the system. Four other numbers matter more.
- Maximum drawdown. The worst peak-to-trough fall. You will experience this. If the number would make you abandon the strategy, you do not have this strategy.
- Longest losing streak, and longest flat period. A system can be profitable overall and spend fourteen months going nowhere. That is where people quit.
- Number of trades. Forty trades is not a sample. Two hundred begins to be. A spectacular result from twelve trades is a story, not evidence.
- Dependence on outliers. Remove the best three trades. If the strategy is now unprofitable, its entire result rests on catching a few specific moves — which may not recur.
The honest final step
Before risking real money, trade the system on paper — or at genuinely tiny size — for a few months. Not to re-validate the maths, which the backtest already covered, but to find out whether you can actually follow the rules when the losses are live and the screen is red. A great many strategies that test well are abandoned in week three by the person running them, and that failure is invisible in any backtest.
Your strategy returns 34% CAGR over 2019–2024 on the current NIFTY 50 constituents, with 41 trades and a 9% maximum drawdown. What is the biggest problem?
Kal ka match dekh ke aap bata sakte ho ki kis over pe kya karna tha — sab saaf dikhta hai. Par live match mein woh nahi pata hota. Backtest mein yahi dhokha hota hai: aap anjaane mein aage ka result jaante ho aur rules usi hisaab se fit kar dete ho. Isko honest rakhna asli skill hai.
- The question is not "did it work" but "would I have chosen these rules before seeing the data".
- Overfitting, look-ahead, survivorship, ignored costs and period selection are the five standard lies.
- Walk-forward testing with genuinely untouched data is the strongest single defence.
- Judge drawdown, losing streaks, sample size and outlier dependence — not just CAGR.
- Paper-trade before going live, to test whether you can follow the rules, not whether they work.
Mark it done to track your progress through the curriculum.
Common questions
Short, direct answers to what people ask about this topic.
- survivorship bias in backtesting meaning
- Survivorship bias is testing a strategy on the companies that are in an index today, which quietly excludes every company that was delisted, suspended or collapsed during the test period. Running ten years of history on today’s NIFTY 50 measures only the firms that made it, so the result is optimistic by a large and unknown margin. The fix is point-in-time index constituents — the list as it actually stood on each date.
- using information in a backtest that was not available at the time of the trade is known as
- Look-ahead bias. It creeps in through small details — buying at today’s open using today’s close, or acting on a quarterly figure before the company had filed it, which in India is often weeks after the quarter ended. The defence is to rebuild the test so every decision uses only data that existed at that exact timestamp.
- how many trades does a backtest need to be meaningful
- Forty trades is not a sample; a couple of hundred begins to be one, because below that you cannot separate a genuine edge from a run of luck. A useful companion check is to delete the three best trades and rerun — if the strategy is now unprofitable, its whole result rests on catching a few specific moves rather than on a repeatable edge.
- what is walk forward testing
- Walk-forward testing means splitting your history, designing the rules on the earlier portion, freezing them completely, and then running them once on later data you have never looked at; the fuller version repeats that split several times, rolling the window forward so every stretch of out-of-sample data is judged by rules that were fixed before it. The frozen step is the entire discipline — if you go back and tweak after seeing the out-of-sample result, that data is now in-sample and worthless as a test. Some degradation out of sample is normal; a collapse from excellent to breakeven means you fitted noise.
- why does a strategy work in backtest but fail in live trading
- The usual culprits are overfitting, omitted costs and a flattering test window. Tuning parameters until the equity curve looks beautiful fits the past and predicts nothing; leaving out brokerage, STT, stamp duty, GST, slippage and the bid-ask spread turns a loss-making system into a profitable-looking one, especially at high trade frequency; and a period like 2020 to 2024 contains no sustained bear phase, so almost every long strategy passes it. A fourth reason is invisible in any backtest — the rules are abandoned in week three when the losses are live.