Aureon
← indexframeworkJUL 23, 202616 min read

A Complete Framework for Evaluating Any Strategy Before You Risk Money

Most traders evaluate a strategy by examining one number and asking one question: did it make money? That question is close to useless in isolation. A strategy can produce a profitable backtest for at least five reasons unrelated to possessing an edge: luck, overfitting, a single favorable regime, unmodeled costs, and position sizing that would have ended the account in live trading.

What follows is a framework for asking better questions. Five gates, applied in order, each designed to eliminate a strategy for a different reason. A strategy clearing all five is not guaranteed to work. It has earned the right to be traded small.

The order is deliberate. Each gate costs more to run than the one before it, so weak ideas fail quickly and real effort is reserved for survivors.


Gate 1: Is there an edge at all?

The question: does this system have positive expectancy, measured properly, across a sample large enough to be meaningful?

Expectancy, not return

Begin with expectancy per trade:

Expectancy = (Win% × Avg Win) − (Loss% × Avg Loss)

Total return can be produced by one exceptional winner within a system that is otherwise broken. Expectancy describes what the average trade is worth, which is what you are actually repeating.

Convert it into a comparable figure by expressing it in units of risk:

Expectancy in R = Expectancy per trade / Average loss

If the average loss is $100 and expectancy is $30, the system produces 0.3R per trade. This allows direct comparison between a day-trading system and a swing system, since both are denominated in risk taken rather than in dollars or percentages carrying different meanings.

Sample size

Here is the uncomfortable arithmetic. Trading returns are noisy, and distinguishing a genuine edge from luck requires considerably more data than intuition suggests.

The scale of the problem: the number of trades required to detect an edge grows with the square of the noise-to-signal ratio. A system with a small edge relative to its trade-to-trade variability, which describes essentially every retail strategy, requires hundreds of trades before the result carries meaningful information. A system with 30 trades has, in statistical terms, communicated almost nothing, regardless of how those 30 trades look.

Practical thresholds:

  • Under 100 trades: not evidence. Treat the result as a hypothesis.
  • 100 to 300 trades: suggestive. Sufficient to justify further testing, insufficient to size up on.
  • 300 or more trades across varied conditions: the number begins to mean something.

The phrase "across varied conditions" carries as much weight as the count. Three hundred trades taken within a single 18-month bull market is one observation of one regime repeated three hundred times rather than three hundred independent pieces of evidence.

Gate 1 fails if: expectancy is negative or statistically indistinguishable from zero, or the sample is too small to determine either.


locked · premium

The rest is for members.

You’re reading the free preview. Premium unlocks the full file on every deep-dive:

  • 01the full methodology: costs, sizing, failure modes
  • 02the complete, runnable research notebook
  • 03every resource unlocked the day it ships

the dispatch

Get the next piece of research when it ships. No schedule filler.

want the full notebooks behind articles like this? go premium →