Common Backtesting Mistakes (and How to Avoid False Edges)

Backtests fail when assumptions are generous. If you ignore fees, borrow costs, partial fills, and market impact, you will overstate performance. Survivorship bias and lookahead errors are quieter killers that only surface when live P&L diverges sharply from the research notebook. Without reproducible seeds, environment captures, and clear data provenance, every rerun becomes a different experiment.
TL;DR — Key Takeaways
- ✦Most backtest results overstate live performance by 30-70% due to survivorship bias, look-ahead bias, and unrealistic fill assumptions — consistent with research on backtest overfitting.
- ✦Walk-forward validation is non-negotiable: in-sample optimisation followed by out-of-sample testing on unseen data is the minimum standard.
- ✦Slippage and brokerage fees must be modelled at your actual broker tier — not theoretical minimums — or profitable-looking strategies go negative in live trading.
- ✦The single most common mistake: optimising too many parameters on too little data, creating a curve-fitted strategy that only works on the backtest period.
Red Flags in Research
Watch for these signals before trusting a backtest:
- Equity curves that never show sideways periods or meaningful drawdowns
- Parameter sets tuned to a single symbol or short date range with no validation set
- Trades that assume instant fills at mid-price with zero queue position
- Data with no corporate action handling, bad ticks, or timezone normalization
- Performance jumps that coincide with data vendor changes or missing delisted symbols
Make Backtests More Honest
Model realistic friction: borrow fees on short positions, maker/taker fees for crypto, and slippage that scales with volatility. Use walk-forward testing with unseen periods, and validate on different instruments or venues to avoid memorizing a regime. Enforce order throttles and queue-depth assumptions so your backtest cannot place liquidity that wouldn't realistically execute.
"If a strategy only works in backtests that ignore fees and impact, it doesn't work."
Once live, compare fill prices, slippage, and latency to the backtest assumptions. Close the loop quickly: adjust models, tighten limits, and keep a log of research changes so you can reproduce every decision. Maintain a changelog of data updates, parameter tweaks, and environment versions so you can trace any divergence to a concrete cause.
Want us to build this for you?
Talk to our teamPractical Backtest Hygiene
Habits that keep research credible:
- Freeze and version datasets; rerun baselines whenever the dataset changes
- Use commission/fee schedules that match your broker tier, not marketing numbers
- Stress-test with worse-than-expected slippage and check if the edge survives
- Compare fills to benchmarks like VWAP or arrival price and chart the drift
- Store seeds and package versions in the notebook so results are reproducible
Already Have a Strategy? Let's Automate It.
At Arkalogi, we convert your trading logic into fully automated systems - integrated with your broker, backtested on real market data, and deployed on a live server. You describe your strategy in plain English. We handle everything else. Book a free honest assessment on WhatsApp to chat with us. No sales pitch. Just clarity on what's possible and what infrastructure you need to avoid common failure modes.
This post was written by Leena Shah, a Machine Learning Engineer at Arkalogi.
If you want a custom strategy like this built for your broker, we can help.