When a Sample Is Too Small to Say Anything

The most common analytical error in trading is not miscalculation. It is drawing a firm conclusion from a run of sessions that was never long enough to distinguish a real difference from ordinary variation. The arithmetic is usually correct. The record says what it says. The problem is that a record that short would have said something equally definite even if nothing at all were going on.
Chance Produces Convincing Patterns

A short sequence of outcomes with any moderate hit rate will routinely contain streaks in both directions. Those streaks look like information because they have shape, and the human reading them is not neutral. If the streak is good, the approach is working. If it is bad, something has changed. Both readings feel like observation and both are usually noise.
This is worse when comparing two variations. Run one version of a range rule for a couple of weeks, run another for a couple of weeks, and one of them will come out ahead. It will come out ahead almost regardless of whether there is any real difference, because the comparison has nowhere near enough sessions to separate a small edge from the spread of ordinary results.
Noisier Results Need More of Them

How large a sample needs to be is not a fixed number. It depends on how spread out the results are and how large a difference you are trying to detect. Two things follow from that, and both are unwelcome.
First, an approach with widely varying session results needs a substantially longer record than one with tightly clustered results, for the same confidence. Second, small differences need far more evidence than large ones. Detecting that a variation is dramatically better takes relatively few sessions. Detecting that it is slightly better can take more sessions than you will accumulate in a reasonable time.
Count Sessions, Not Days
A related trap is describing the sample by how long it took rather than how many observations it contains. Six months of a rule that trades rarely may contain fewer usable results than a few weeks of one that trades every session. The calendar is reassuring and irrelevant.
It also matters whether the sessions are genuinely independent. A run of sessions inside one sustained market condition is closer to one observation repeated than to many separate observations, no matter how many rows it fills. A record that looks adequate in length but covers only a single kind of environment has not tested the approach against the conditions it will eventually meet.
Saying Nothing Is a Valid Output
The discipline here is being willing to finish a review with no conclusion. That is genuinely difficult, because a review that produces nothing feels like wasted effort, and because there is usually a decision waiting on the answer.
But an unsupported conclusion is worse than none. It leads to changing a rule that was fine, or scaling up something that has not been demonstrated, and both of those cost more than patience. Writing down what you would need to see, and how many more sessions it would take, converts the non answer into something actionable without inventing certainty.
Working Sensibly While You Wait
Waiting does not mean doing nothing. It means changing one thing at a time so that the record accumulating is about a single question rather than several tangled ones. It means keeping the session level detail so that a proper look is possible later. And it means resisting the urge to re-examine the numbers every few days, since checking repeatedly and stopping when the answer looks good is a reliable way to be convinced of something that is not there.
Decide in advance roughly how many sessions you intend to gather before looking properly, then look then. The result will be more honest than any number of interim glances, and the interim glances were never going to tell you anything the final look would not.