A Complete Framework for Evaluating Any Strategy Before You Risk Money
Most traders evaluate a strategy by examining one number and asking one question: did it make money? That question is close to useless in isolation. A strategy can produce a profitable backtest for at least five reasons unrelated to possessing an edge: luck, overfitting, a single favorable regime, unmodeled costs, and position sizing that would have ended the account in live trading.
What follows is a framework for asking better questions. Five gates, applied in order, each designed to eliminate a strategy for a different reason. A strategy clearing all five is not guaranteed to work. It has earned the right to be traded small.
The order is deliberate. Each gate costs more to run than the one before it, so weak ideas fail quickly and real effort is reserved for survivors.
Gate 1: Is there an edge at all?
The question: does this system have positive expectancy, measured properly, across a sample large enough to be meaningful?
Expectancy, not return
Begin with expectancy per trade:
Expectancy = (Win% × Avg Win) − (Loss% × Avg Loss)
Total return can be produced by one exceptional winner within a system that is otherwise broken. Expectancy describes what the average trade is worth, which is what you are actually repeating.
Convert it into a comparable figure by expressing it in units of risk:
Expectancy in R = Expectancy per trade / Average loss
If the average loss is $100 and expectancy is $30, the system produces 0.3R per trade. This allows direct comparison between a day-trading system and a swing system, since both are denominated in risk taken rather than in dollars or percentages carrying different meanings.
Sample size
Here is the uncomfortable arithmetic. Trading returns are noisy, and distinguishing a genuine edge from luck requires considerably more data than intuition suggests.
The scale of the problem: the number of trades required to detect an edge grows with the square of the noise-to-signal ratio. A system with a small edge relative to its trade-to-trade variability, which describes essentially every retail strategy, requires hundreds of trades before the result carries meaningful information. A system with 30 trades has, in statistical terms, communicated almost nothing, regardless of how those 30 trades look.
Practical thresholds:
- Under 100 trades: not evidence. Treat the result as a hypothesis.
- 100 to 300 trades: suggestive. Sufficient to justify further testing, insufficient to size up on.
- 300 or more trades across varied conditions: the number begins to mean something.
The phrase "across varied conditions" carries as much weight as the count. Three hundred trades taken within a single 18-month bull market is one observation of one regime repeated three hundred times rather than three hundred independent pieces of evidence.
Gate 1 fails if: expectancy is negative or statistically indistinguishable from zero, or the sample is too small to determine either.
Gate 2: Is it robust, or did you fit it?
The question: does the edge survive changes that should not matter?
A genuine effect occupies a broad region of parameter space. An overfit result occupies a single point. The test is determining which one you hold.
The parameter grid
Take every parameter and re-run the strategy across a range surrounding the chosen value, then examine the entire surface.
A real edge produces positive performance across a wide contiguous region, degrading gradually as parameters move away from the center. The chosen parameters are not dramatically superior to their neighbors. They sit in the middle of a plateau.
Overfitting produces a sharp spike. The exact parameters work and one step in any direction collapses the result. You have located the specific configuration that fit this history's noise.
A practical rule: prefer the center of a broad plateau to the peak of a narrow spike, even though the plateau's reported numbers are worse. The spike's superior backtest performance is precisely the component that will not repeat. Selecting the plateau costs backtest performance and purchases live performance.
Out-of-sample, once
Develop on one period, test on another you have never examined, and run that test a single time. Every additional review-and-adjust cycle contaminates the holdout and converts your only unbiased estimate into further fitted data.
For a stronger version, use walk-forward analysis: optimize on a rolling window, test on the period immediately following, roll forward, and repeat. This approximates how a system would be retuned in real time and produces a series of genuinely out-of-sample results rather than one.
The degradation ratio
Compare out-of-sample performance to in-sample directly:
Degradation = Out-of-sample Sharpe / In-sample Sharpe
Some decay is normal, since an in-sample result always contains fitting. As a rough guide, retaining most of the in-sample performance is a strong sign, retaining roughly half is typical of a real but modest edge, and retaining little or none indicates fitted noise. A sign flip means the strategy is worse than doing nothing.
Gate 2 fails if: performance forms a spike rather than a plateau, or out-of-sample results collapse relative to in-sample.
Gate 3: Which regime does this require?
The question: is this an edge, or a bet on one market environment that happened to dominate the sample?
A large number of strategies that appear to work are a single implicit bet, long volatility, short volatility, long equity beta, or momentum-friendly conditions, and they work exactly as long as that condition persists.
Slice the results
Break performance down along every available dimension.
By year. Is the total driven by one exceptional year? Remove the single best year and check whether the strategy remains profitable. If not, you hold one trade rather than a strategy.
By volatility regime. Split into high and low volatility periods. Many mean-reversion systems exist entirely within high-volatility windows, and many trend systems fail there.
By market direction. Bull, bear, and sideways. A long-biased system that generated everything during a bull market has described the market rather than itself.
By best trade. Remove it. If the strategy is now unprofitable, the result was one favorable outcome presented as a system.
Correlation to the obvious
Compare returns to a simple benchmark, either buy and hold or a basic momentum rule. If correlation is high and risk-adjusted performance is similar, you have constructed a complicated method of doing a simple thing, and the complexity is pure added fragility.
The most valuable strategies produce return streams that do not resemble anything else you hold. That property is also what makes them worth adding to a portfolio.
Gate 3 fails if: returns collapse on removal of one year or one trade, or the strategy is an undisclosed bet on a regime you cannot identify in advance.
Gate 4: Does it survive real costs?
The question: how much gross edge remains after the market takes its share, and how much room exists before the edge disappears entirely?
Model every cost
- Spread: paid on entry and exit, every trade.
- Commissions: per trade or per share.
- Slippage: the gap between expected and realized fill. It is not constant, widening precisely when volatility spikes, which is when many strategies want to trade.
- Market impact: at size, your own order moves the price against you.
- Financing and borrow: for leverage or short positions.
The breakeven cost test
This is the most informative single figure in the gate. Rather than asking whether the strategy is profitable at 5 basis points, ask:
At what cost per trade does expectancy reach zero?
Compare that to your realistic execution cost. The gap between them is the margin of safety.
Breakeven cost far above realistic cost: durable. Execution quality is not the deciding variable.
Breakeven cost near realistic cost: fragile. Your broker, fill quality, and volatile-day slippage now determine profitability rather than the strategy.
Breakeven cost below realistic cost: the edge exists only on paper.
Turnover is the multiplier
Cost drag scales with trading frequency:
Annual cost drag ≈ round trips per year × cost per round trip
A strategy trading daily pays that cost hundreds of times annually. One trading monthly pays it twelve times. This explains why high-frequency retail strategies fail live so often despite acceptable backtests. The gross edge was real and turnover consumed it. When comparing two strategies with similar gross performance, the lower-turnover option is almost always the better bet, because it retains more room before costs become decisive.
Gate 4 fails if: breakeven cost is close to or below realistic execution cost.
Gate 5: Can you survive trading it?
The question: at the size you would trade this, does a normal adverse stretch leave you solvent, and would you still be holding it?
A strategy can clear every previous gate and still end the account, because expectancy assumes a large number of trades and position sizing determines whether you take them.
Size against the worst plausible streak
Losing streaks are arithmetic rather than misfortune. In a system winning half the time, five consecutive losses appear roughly once in every thirty-two sequences, a near certainty across a year of active trading. The maximum losing streak observed in a backtest is a lower bound. The future will eventually produce a longer one.
Size so that the worst plausible streak, meaningfully longer than the worst in your data, is survivable rather than terminal. Then check the recovery arithmetic, since losses and gains are not symmetric:
Down 10% → need +11% Down 50% → need +100%
Down 20% → need +25% Down 67% → need +200%
A drawdown requiring the account to triple in order to recover is not a setback. Sizing conservatively is not caution for its own sake. It keeps the recovery arithmetic within a range where recovery remains realistic.
The drawdown you can hold
A backtest's maximum drawdown is a figure on a screen. Living through it is different. It arrives slowly, accompanied by the persistent question of whether the edge has stopped working, and it typically precedes the gains.
Two questions worth answering honestly.
Would you hold through this drawdown, in real money, without knowing when it ends? If not, the strategy is unusable at that size regardless of its Sharpe ratio. A strategy abandoned at the bottom has negative expectancy in practice, whatever the backtest reported.
Would you distinguish a normal drawdown from a broken strategy? Define, in advance and in writing, what would constitute evidence that the edge has stopped working: a drawdown beyond a specified threshold, a losing streak longer than any observed in testing, or a change in the market structure the hypothesis depended on. Deciding this before the pain arrives is the only point at which you can decide it clearly.
This is what having numbers your psychology can rely on means in practice. Conviction fails under pressure. Evidence gathered in advance is what remains.
Gate 5 fails if: the drawdown at intended size is not survivable, or you cannot articulate in advance what would signal a stop.
The scorecard
| Gate | Test | Pass condition |
|---|---|---|
| 1. Edge | Expectancy in R, sample size | Positive expectancy, 300+ trades across conditions |
| 2. Robustness | Parameter grid, out-of-sample | Broad plateau; OOS retains meaningful performance |
| 3. Regime | Year, volatility, and direction slices; drop best trade | Survives removal of best year and best trade |
| 4. Costs | Breakeven cost vs. realistic cost | Comfortable margin between them |
| 5. Survivability | Streak sizing, drawdown tolerance | Worst plausible streak survivable; stop rule defined |
A strategy failing any gate receives no capital. Not reduced size. No capital, until you understand why it failed and whether the failure is correctable without fitting.
A strategy clearing all five is traded small, with real money, for long enough to compare live results against backtested expectations. Live trading is the final out-of-sample test and the only one whose data cannot be contaminated by prior examination.
Why this order
The gates run from cheapest to most expensive by design. Gate 1 takes minutes and eliminates most ideas. Gate 2 takes hours. Gates 3 and 4 require substantial work. Gate 5 requires honesty about yourself, the scarcest input available. Failing quickly at the top means limited research time is spent on the small number of hypotheses that warrant it.
The framework's real function is not identifying winners. It is making self-deception difficult, which, given that the person selling you on a strategy is usually you, is the more valuable capability.
Next: deconstructing the debit spread, covering the specific conditions under which the structure's arithmetic genuinely favors you and how to test it systematically.
put it to work
The Backtest Audit Checklist turns this into something you can run against your own strategy. Free, PDF, no card needed.
get it freeevery tool behind the research lives in resources →
