How to Read a Backtest: 8 Checks Before You Trust It
A backtest replays a strategy's rules over past prices to show what would have happened. This guide walks through eight checks — costs, trade count, drawdown, market regimes, in-sample tuning, overfitting, cherry-picking and position sizing — and ends with a checklist table you can reuse on any strategy, including ours.
A backtest is a model, not a record: its equity curve is only as honest as its assumptions.
1. Costs: fees, slippage and funding
Check costs first, because they repeat on every trade.
- Fees are charged on entry and on exit. A test should apply them to both sides of every trade.
- Slippage is the gap between the price the rules signal and the price you actually get. It grows in fast markets, on thin order books and with market orders.
- Funding applies to perpetual futures: longs and shorts pay each other at set intervals while a position is open, so multi-day holds can pay it many times. How it works is covered in liquidation, margin and funding.
A hypothetical example shows why this matters. Assume a 0.05% fee per side plus 0.05% slippage per side, or 0.2% per round trip. On a 1,000 USDT position that is 2 USDT per trade, so a strategy that trades 200 times a year spends 400 USDT on costs before a single winner counts. A result shown "before costs" can turn negative once costs are added, and strategies that trade often are hit hardest.
2. Sample size: count the trades
Twenty trades can look brilliant by luck; a few hundred trades across different conditions say much more. There is no magic threshold, but the fewer the trades, the more a single outlier decides the result.
Three quick tests:
- Remove the best two or three trades. If the result collapses, it rests on outliers.
- Find the longest losing streak. You would have to live through it with real money.
- Ask for the full trade list, not only the summary. Anyone can publish a curve; a list of every entry and exit can be checked against the chart.
3. Drawdown and the market regimes covered
Maximum drawdown is the largest fall from an equity peak to a later low. It shows how painful the strategy was to hold, which often matters more than the final result. Check what capital it is measured against and how long the account stayed below its previous peak. More in max drawdown explained.
Then check the period. A test that covers only a rising market says little about a falling or sideways one. Look for at least one strong uptrend, one steep decline and one long range. If results by regime aren't shown, scroll the trade list against the price chart yourself; bull, sideways and bear markets explains what to look for.
4. In-sample tuning, overfitting and cherry-picking
In-sample data is the period used to design and tune the rules. Out-of-sample data was not used for tuning: a later period, or other markets. A result measured only in-sample is close to a best case, because the settings were chosen for working on that data. Expect live results to be worse.
Overfitting means tuning so closely to past noise that the rules stop working on new data. Signs to watch for:
- Oddly precise parameters (a 37-bar average, a 2.35 multiplier) where nearby values give very different results.
- Many filters and exceptions, each added to fix one past losing trade.
- An equity curve that is almost a straight line.
A simple robustness check is to nudge each parameter a little. A sturdy strategy degrades gradually; an overfit one falls apart.
Selection bias is cherry-picking: showing the coin, timeframe or period where the rules happened to work. Ask how many symbols and timeframes were tested, and where the rules failed. Also watch for survivorship bias: coins that collapsed or were delisted don't appear in today's lists, so tests run only on today's top coins look better than reality.
5. Position sizing: what capital are the numbers based on?
The same trades produce very different curves depending on sizing. Check three things:
- Allocation per trade: a fixed amount, or a share of equity that compounds?
- Leverage and add-ons: adding to a position multiplies exposure, and with leverage a drawdown can exceed the amount allocated.
- Capital base: is the drawdown a share of the whole account or of the slice allocated to the strategy? The same loss looks small against one and large against the other.
Also note what a backtest cannot model: liquidation, margin calls, exchange outages, partial fills and your own hesitation.
6. A backtest checklist you can reuse
| Check | What to look for | Red flag |
|---|---|---|
| Costs | Fees on every entry and exit, slippage, funding for perpetuals | "Gross" results, costs not stated |
| Trade count | Hundreds of trades, full trade list | A handful of trades, summary only |
| Max drawdown | Stated against a clear capital base, with recovery time | Only returns shown |
| Regimes | Uptrend, decline and range all covered | One trending period |
| In-sample vs out-of-sample | Which data was used for tuning | Not mentioned |
| Selection | All symbols and timeframes tested are disclosed | Only the best coin shown |
| Sizing | Allocation, leverage, compounding and add-ons explained | Unclear capital base |
| Overfitting | Similar results with nearby settings | Oddly precise parameters, near-perfect curve |
How Pullwave publishes its backtests
Pullwave publishes every backtest and every trade on its performance page, so you can run this checklist yourself. Fees are included in all of them, and the futures backtests also include funding and slippage. The spot backtests don't include slippage, so read them as somewhat optimistic, and the futures backtests don't model liquidation or margin. The settings were also refined on the same past data, which makes these in-sample results: future results could be worse.
FAQ
Q. What is the difference between in-sample and out-of-sample results?
In-sample results come from the same data used to design and tune the rules. Out-of-sample results come from data the rules never saw during tuning. They are the more realistic preview, and they are usually weaker.
Q. How many trades does a backtest need?
There is no fixed number, but more trades across more market conditions make luck a weaker explanation. Be cautious with a few dozen trades, and check whether removing the best few changes the picture.
Q. Why do live results often come out worse than the backtest?
Real fills include slippage and delays, tuned settings meet new data, and markets move into conditions the test never contained. Traders also skip signals or change size after losses, which a backtest never does.
Q. Is a smooth equity curve a good sign?
Not by itself. Real strategies have losing streaks and drawdowns, so a near-perfect curve more often points to overfitting, missing costs or a short, favorable period. Treat it as a reason to check harder.
Note: This article is general education, not investment advice. Backtest results describe the past only and do not guarantee future results. Futures can be liquidated and you can lose your capital. Start small.