ORB Trading Statistics

Which figures are worth computing from a run of sessions, and how to read them without flattering yourself. Distribution ahead of average, the influence of a single outlier day, and knowing when a sample cannot support a claim.

A Run of Sessions Is Data, If You Treat It Like Data

Anyone trading the open accumulates a record whether they intend to or not. Turning that record into something informative requires deciding, in advance, which quantities are worth computing and what each of them can honestly support. Most trading records are summarised by a single number, usually a total or an average, and a single number is the least informative thing a set of sessions can produce. The interesting structure is in the spread, the shape and the exceptions.

Averages Hide the Sessions That Matter

An average result per session tells you almost nothing about what a session feels like or what range of outcomes to prepare for. Two records with identical averages can be built from completely different distributions, one clustering tightly around the middle and the other consisting largely of small losses redeemed by occasional large wins. Those two records demand different position sizing, different psychological preparation and different amounts of capital to survive, and the average is silent on all of it.

Outliers Deserve Individual Attention

In a modest sample, one exceptional session can dominate the total. That single result then quietly determines whether the record looks profitable, and every conclusion drawn from the record is really a conclusion about that one day. Recomputing the figures with the largest contributor removed is not a way of hiding an inconvenient truth. It is a way of finding out how much of the story depends on something that happened once and may not happen again soon.

Small Samples Should Be Allowed to Say Nothing

The strongest temptation in reviewing a trading record is to extract a conclusion from it regardless of how many sessions it contains. A handful of sessions can produce any result at all through ordinary chance, and reading a preference out of that is how a working approach gets abandoned and a poor one gets scaled up. Knowing roughly how large a sample needs to be before a difference means anything is more valuable than any individual statistic.

Numbers Worth Computing

The articles here deal with which figures to compute from a series of sessions and how to read them without flattering yourself. They cover reading a distribution rather than its centre, identifying and testing the influence of outlier sessions, and recognising when a sample is simply too small to support a claim. They do not recommend entries, stops or range lengths. They are about measuring whatever you already do.

Latest Guides

Close-up of a live cryptocurrency trading chart screen displaying dynamic market trends and analysis.

Read the Distribution, Not the Average

2026-09-03

The average is the first number anyone computes and the last one that should be trusted on its own. It compresses a run of sessions into a single value by design, and the thing it discards is precisely the thing you need in order to know what to expect tomorrow. Two records with the same average can require completely different behaviour from the person trading them.

The Same Centre, Two Different Approaches

Monitors displaying stock market charts in a dimly lit room, perfect for finance and trading themes.

Consider two records over the same number of sessions with identical averages. In the first, almost every session lands close to the middle, with modest wins and modest losses and no result far from the rest. In the second, most sessions are small losses and the total is carried by a handful of large wins.

The first record is comfortable to trade and can be assessed relatively quickly, because each session contributes similar information. The second requires the temperament to sit through long unprofitable stretches and enough capital to still be present when the rare large session arrives. Nothing in the average distinguishes them, and everything about how you would run them does.

What to Look At Instead

Detailed financial trading screen with colorful charts and data representing market fluctuations.

The most useful first step is simply to sort the sessions and look at them in order, from worst to best. This costs nothing and immediately shows the shape. You see whether the losses are tightly grouped or ragged, whether the wins taper smoothly or jump, and whether there is a gap somewhere in the middle where results are strangely absent.

Beyond that, the middle value is worth having alongside the average. When the two sit close together the record is reasonably symmetric. When the average sits well above the middle, the total is being pulled by the upper tail, and the typical session is worse than the headline suggests. That single comparison catches most of the cases where an average is misleading.

The Tails Are the Operational Question

The best sessions determine whether the approach is worth running. The worst sessions determine whether you can run it. These are different questions and only the second one can end the experiment early.

So the left tail deserves specific attention. How bad was the worst session, how many sessions like that occurred, and did they cluster together or arrive spread out. A record with an acceptable average and a cluster of severe losses in one week describes an approach that could remove you from the market before its average has a chance to assert itself.

Dispersion Deserves a Number

Once you have looked at the shape, it helps to have a rough measure of how spread out the results are, so that comparisons across periods are possible. The specific measure matters less than using the same one consistently. A simple spread between the typical worst and typical best session is enough for most purposes and is easier to reason about than a formal statistic.

The value of having this recorded is that it makes changes visible. If the average holds steady but the spread widens, the approach is producing the same total from wilder swings, which is a real change in what you are trading even though the headline figure did not move.

Keep the Session Level Record

None of this is possible after the fact if the record only stores totals by week or by month. Aggregation destroys the distribution and it cannot be recovered later. The single most valuable habit in this area is keeping one row per session with the outcome on it, in whatever form is convenient, for long enough that the run becomes worth examining.

Everything else can be recomputed from that. Averages, middles, spreads, tails and any comparison you decide later that you want to make all derive from the same list. A summary computed today and stored instead of the underlying sessions answers exactly one question, and it will not be the question you have in six months.

Read more →

A man celebrates success at a multi-monitor workstation while analyzing stock charts.

The One Outlier Session That Carries Your Whole Result

2026-09-03

There is a specific and uncomfortable exercise worth performing on any trading record. Find the single best session in the run, remove it, and recompute everything. If the record was profitable before and is not profitable afterwards, then the conclusion you have been carrying around is not really about your approach. It is about one day.

Why This Happens So Easily

A multi-monitor stock trading setup showcasing charts and data analysis in a home office setting.

Breakout approaches tend to produce asymmetric results by construction. Losses are bounded by wherever the stop sits, while the occasional session that runs a long way is not bounded by anything except when you decide to exit. Over a modest number of sessions, that asymmetry means one or two results will usually be far larger than the rest.

This is not a flaw. It is the mechanism the approach relies on, and a record without any large sessions would be a different kind of warning. The flaw is in treating a total that depends heavily on those sessions as though it were a stable estimate of what the approach produces per day.

The Leave One Out Check

A businessman at his desk analyzing financial charts on multiple monitors.

Removing the largest contributor and recomputing is the crudest version of a robustness check and it is enough to be informative. Do it for the best session and separately for the worst. Then do it for the two largest of each.

What you are looking for is not whether the numbers change. They will. You are looking for whether the sign changes, or whether a comparison between two variations of your approach reverses. A conclusion that survives having its biggest contributor deleted is a conclusion worth acting on. A conclusion that flips is a statement about one session dressed up as a statement about a method.

Removal Is a Diagnostic, Not a Correction

It is important to be clear about what this exercise does not license. The large session happened. It was a real trade, taken under the real rules, and it belongs in the record permanently. Deleting outliers to produce a tidier looking result is exactly the wrong lesson.

The purpose is to learn how much of your confidence rests on a single event. If the answer is most of it, that does not mean the approach is bad. It means the record has not yet accumulated enough of the sessions that make the approach work, and the correct response is to keep going and check again later rather than to conclude anything now.

Ask What Kind of Day It Was

The other half of the work is qualitative. An outlier session usually has a cause, and the cause determines whether more of them should be expected. A session that ran a long way because a scheduled event repriced the instrument is a different thing from a session that ran a long way because a quiet range broke into a sustained trend.

If the outlier arrived on a day your rules would now exclude, then the record contains a profit your current rules could not have earned, and the record overstates what the current version of the approach would do. That is a common and easily missed situation, because rules evolve while the record stays fixed.

Outliers on the Other Side

The same discipline applies to the worst session, and it is applied far less often, because a large loss tends to be explained rather than examined. It gets attributed to a mistake, a technical problem, or an unusual day, and then mentally set aside as unrepresentative.

Sometimes that is right. Often it is not, and the same conditions will recur. A single severe loss that is dismissed as exceptional and then repeats twice more over the following months was never exceptional. It was the tail of the distribution making its first appearance, and treating it as an outlier meant the record was quietly understating what the approach can cost. The test for both sides is the same. Remove it, recompute, and see how much of what you believe was resting on it.

Read more →

Financial chart displayed on monitor showcasing stock market trends and analysis.

When a Sample Is Too Small to Say Anything

2026-09-03

The most common analytical error in trading is not miscalculation. It is drawing a firm conclusion from a run of sessions that was never long enough to distinguish a real difference from ordinary variation. The arithmetic is usually correct. The record says what it says. The problem is that a record that short would have said something equally definite even if nothing at all were going on.

Chance Produces Convincing Patterns

Close-up of a live cryptocurrency trading chart screen displaying dynamic market trends and analysis.

A short sequence of outcomes with any moderate hit rate will routinely contain streaks in both directions. Those streaks look like information because they have shape, and the human reading them is not neutral. If the streak is good, the approach is working. If it is bad, something has changed. Both readings feel like observation and both are usually noise.

This is worse when comparing two variations. Run one version of a range rule for a couple of weeks, run another for a couple of weeks, and one of them will come out ahead. It will come out ahead almost regardless of whether there is any real difference, because the comparison has nowhere near enough sessions to separate a small edge from the spread of ordinary results.

Noisier Results Need More of Them

Monitors displaying stock market charts in a dimly lit room, perfect for finance and trading themes.

How large a sample needs to be is not a fixed number. It depends on how spread out the results are and how large a difference you are trying to detect. Two things follow from that, and both are unwelcome.

First, an approach with widely varying session results needs a substantially longer record than one with tightly clustered results, for the same confidence. Second, small differences need far more evidence than large ones. Detecting that a variation is dramatically better takes relatively few sessions. Detecting that it is slightly better can take more sessions than you will accumulate in a reasonable time.

Count Sessions, Not Days

A related trap is describing the sample by how long it took rather than how many observations it contains. Six months of a rule that trades rarely may contain fewer usable results than a few weeks of one that trades every session. The calendar is reassuring and irrelevant.

It also matters whether the sessions are genuinely independent. A run of sessions inside one sustained market condition is closer to one observation repeated than to many separate observations, no matter how many rows it fills. A record that looks adequate in length but covers only a single kind of environment has not tested the approach against the conditions it will eventually meet.

Saying Nothing Is a Valid Output

The discipline here is being willing to finish a review with no conclusion. That is genuinely difficult, because a review that produces nothing feels like wasted effort, and because there is usually a decision waiting on the answer.

But an unsupported conclusion is worse than none. It leads to changing a rule that was fine, or scaling up something that has not been demonstrated, and both of those cost more than patience. Writing down what you would need to see, and how many more sessions it would take, converts the non answer into something actionable without inventing certainty.

Working Sensibly While You Wait

Waiting does not mean doing nothing. It means changing one thing at a time so that the record accumulating is about a single question rather than several tangled ones. It means keeping the session level detail so that a proper look is possible later. And it means resisting the urge to re-examine the numbers every few days, since checking repeatedly and stopping when the answer looks good is a reliable way to be convinced of something that is not there.

Decide in advance roughly how many sessions you intend to gather before looking properly, then look then. The result will be more honest than any number of interim glances, and the interim glances were never going to tell you anything the final look would not.

Read more →