Vulcan Markets
Published 14 September 2026
A concept, explained

How many years does a "best day of the month" need before you should believe it?

Where you are. Someone posted that a particular day of the month is the best day to be long, with a bar chart to prove it. You are about to plan a week around it.

A seasonal average is a mean over a set of observations. Before you know how many observations built it, you do not have a pattern - you have a number that looks like one. The sample size is not the footnote. It is the unit of measurement.

This matters because most seasonality claims circulate without it. "The third Friday of the month is bullish." "End-of-month sessions run higher." Each of those is a claim about an average. The average might be real. It might be three data points. You cannot tell from the headline, and the headline is usually all you get.


What a sample size tells you that a mean does not

Every average has a distribution underneath it. A large sample compresses that distribution - the error bar around the mean tightens, and the mean becomes a more reliable description of what actually tends to happen. A small sample leaves the distribution wide open. The error bar stays broad, which means the mean could shift substantially with only a few more observations.

The error bar is not a hedge. It is specific information. A wide error bar tells you the data does not yet resolve to a reliable pattern at this cell. A narrow error bar tells you the pattern is stable across the sample. Both are actionable reads - one toward patience, one toward attention. Collapsing them into a single colour on a heat map discards the most important dimension.

The correct way to read a seasonality cell is to look at three numbers together: the directional average, the sample count, and the error bar. Read only the average and you are reading less than half the information.


Two dates can share the same average and deserve different confidence

Here is the logic that matters most when you are comparing cells.

Take two dates with identical directional averages. One is built from a large pool of observations - sessions spread across decades, consistent enough that the error bar is narrow. The other is built from a small pool - fewer historical instances, the kind of cell that appears when you cut the data along multiple dimensions simultaneously. Same number in the mean column. Completely different reliability.

When you read only the mean, you treat them as equivalent. They are not. The first has been tested more times and survived. The second is a hypothesis that has not accumulated enough evidence to confirm or deny itself. Weighting them identically is a method error, not a trading decision.

The discipline is to see the difference before you form a view.


Why more dimensions shrink every cell

A seasonality study that analyses 30 years of sessions by calendar day alone gives each day a large, coherent pool of observations. Add a second dimension - say, session behaviour by ATR regime - and you have split the pool. Add a third - gap behaviour - and you have split it again. Each additional cut produces more granular cells and smaller samples inside each one.

This is not a flaw. It is the correct constraint to surface. More dimensions mean more specificity, and specificity has a cost in sample depth. A well-constructed seasonality surface shows you both: the more precise read and the smaller pool that produced it. The error bar widens as the cell's sample shrinks. That widening is the surface telling you something true.

The problem arises when the surface hides the constraint - when it shows you only the directional average and lets you infer a precision the data does not support. Showing the sample size and error bar alongside every cell is what prevents that inference. It limits what the display implies to what the statistics support.

The same logic runs through pair-trading research. When a cointegration test fails, the correct response is to say so and cap the implied confidence accordingly - not to present the output as though the test had passed. The constraint is the information.


Read the cell, not just the colour

The practice that separates a research tool from a heat map is whether it shows you all three numbers. A directional average coloured green is not evidence. A directional average coloured green, with a sample count that supports it and an error bar narrow enough to make the average meaningful - that is closer to evidence.

Error Learning on Vulcan Trading surfaces 30 years of S&P sessions by calendar day across more than 14 breakdown dimensions. Every cell carries its sample size and error bar alongside the directional average. The methodology does not hide the constraint when the sample is thin. It displays it.

That is the discipline. When you read a seasonality cell, read all three numbers. The average tells you what tended to happen. The sample size tells you how many times. The error bar tells you how much to trust the average as a description of what will tend to happen next. None of the three is optional. Start free at vulcan-trading.ai.

The Election cycle study in Vulcan Trading's Error Learning surface: a verdict box naming seven election years in the history, a chart of the year paths with their min-max and interquartile bands, and the caption that the spread is the finding.
Where this lands in Vulcan Trading. A seasonal claim with its sample size in the open: seven observations, a spread wider than the average, and a verdict written in words.
  1. The study states its sample in words before its chart: election years come every four years, so the history holds seven of them, and two were bear markets that move the average more than any seasonal effect.
  2. The shaded band is the full min-max range of the underlying years, the darker band the interquartile range. The spread is the finding.
  3. The surface says it in its own caption: where the band holds both large gains and large losses, the average describes arithmetic rather than an expectation.

The objection

You might say
The average is clearly positive. What does the sample size add?
The answer
It tells you how much of that average is signal and how much is a few lucky years. Two cells can share an average and deserve different confidence: one built on decades of sessions, the other on a handful. The surface shows the count and the error bar on every cell so you can read the difference before you act on it.

How Vulcan computes this

Surface
Error Learning
What it does
Thirty years of index sessions by calendar day, across fourteen or more breakdowns, every cell carrying its sample size and its error bar.
Instruments
S&P 500, thirty years
Method
Every cell carries its sample size and error bar, so a result resting on a handful of sessions cannot pass for a strong one. The surface states that it is read alongside a signal, never as one.

Described as capability. This page reproduces no figure from the product and nothing on it is a live read; the method is the point, not a result.

First action

Open Error Learning, pick the calendar-day study, and hover a cell you would have traded: read its count and error bar before its average.

See it on the surface, free. One free account opens the Retail view: the desk, the watchlist, Stocks with the pairs lab inside it, Gold Vault, Order Flow, Forecast and Replay. No card. You are leaving a page about method for a product that applies it. The page stays here.

Create a free account

Next in this path

Door A seasonal pattern burned you · page 1 of 3

Does the market really have good and bad days of the year?Calendar-day seasonality - what three decades of sessions can and cannot say

Door You are sizing a rule and want to know when it stops working · page 3 of 3

That is the end of this path. Pick another door, or read on below.