How significance works

Insights decides whether a difference between A/B test variations is real, or just the noise you would expect from splitting traffic. This page explains the rules behind that decision, so you know what a verdict means and why a test that looks decided might not be.

note

Significance applies to A/B tests only. Personalizations get uplift figures but no winner, because their variations target different audiences rather than competing for the same one, so there is no fair comparison to declare.

Insights uses a two-sided two-proportion z-test at 95% confidence.

  • Two-proportion because it compares two conversion rates: the control's and the variation's.
  • Two-sided because it detects a difference in either direction. A variation that is significantly worse than control is as real a finding as one that is better, and Insights reports it.
  • 95% confidence means a result is called significant when there is less than a 5% probability of seeing a difference at least this large if the two variations were truly identical.

The standard error uses the pooled conversion rate across both arms. This matters for low-traffic tests: without pooling, an arm with zero conversions would have zero variance and drop out of the calculation entirely, overstating how much the data actually tells you.

A trial is one session, not one page view. A visitor who views the same variation three times in one session is one chance to convert, not three. Counting views would treat correlated behaviour from a single person as independent evidence and make results look more certain than they are.

For the same reason, a visitor who converts twice in one session counts once.

A z-score is only meaningful on a sample the test can actually be applied to. Insights checks three things first, and reports no verdict until all of them hold.

RequirementWhy
The counts have to be possibleMore conversions than sessions, or a negative count, means something is wrong with the data rather than with the variation
At least 200 sessions in both the control and the variationA flat floor below which any result is too fragile to act on
At least 5 expected conversions and 5 expected non-conversions per arm, at the pooled rateThe conventional floor for the normal approximation the z-test relies on. Below it, neither the p-value nor the uplift means anything: one conversion in one arm and none in the other reads as a -100% uplift on paper

The second and third are different problems with different fixes. Falling short of 200 sessions means you need more traffic. Clearing 200 sessions but not the expected-count floor means you have plenty of traffic and almost no conversions, so you need a goal that fires more often, or a bigger effect, rather than more visitors.

Conversion rate drives everything. Insights also runs a power analysis, calculating how many sessions per arm you would need to reliably detect a 10% relative improvement at 80% power, and the answer scales harshly as the baseline rate falls. A 10% baseline conversion rate needs roughly 14,700 sessions per arm. A 2% baseline needs roughly 80,700.

That number describes how long to run a test, not whether the test that has already run says anything. Insights deliberately does not use it to gate a verdict. A large, clear effect is significant regardless of whether the sample reached the planned size. Falling short only means a null result cannot rule out a smaller effect, which is why the explanation for "no difference found" mentions being underpowered and the explanation for a winner does not.

When Insights cannot call a result, it says which of these is the reason:

ReasonWhat it means
Not enough dataOne or both arms are below 200 sessions. Keep running
Not enough conversionsEnough sessions, too few conversions among them for a rate to be believable
Not statistically significantThe rates differ, but not by more than noise explains
Identical ratesThe variations are performing exactly the same so far
Invalid dataThe counts are impossible. Worth investigating rather than waiting out

A non-control variation wins by being statistically significant and better than control on the selected metric. For bounce rate this is inverted (lower is better), and Insights accounts for that.

The control can also win. If any variation is significantly worse than control, the control is the winner, and continuing to serve that variation is costing you.

The significance indicator in the dashboard's active tests table is stricter than the winner shown in the composition editor's Insights tab. Both share the same z-test and the same sample requirements, but the dashboard adds two things:

Practical significance. A result also has to clear a 10% minimum detectable effect to count. A 2% uplift measured across enormous traffic can be statistically airtight and still not worth acting on. The dashboard treats it as no result; the Insights tab will still call it a winner.

A Bonferroni correction. When a test has more than one non-control variation, the dashboard divides the 5% error budget across the comparisons, so a two-variation test needs 97.5% confidence rather than 95%. Testing several variations at once gives noise more chances to look like a finding, and this compensates.

The dashboard also pools all non-control variations into a single arm and reports one verdict for the whole test, whereas the Insights tab evaluates each variation against control separately.

The practical consequence is that a test can show a trophy in the Insights tab and a "more data needed" clock on the dashboard at the same time. The two are answering different questions. The tab asks whether the difference is real, and the dashboard asks whether it is real, unlikely to be a multiple-comparison artifact, and large enough to matter.

Hovering the dashboard's significance or uplift indicator shows the p-value, z-score, observed effect, minimum detectable effect, sample sizes, and whether a Bonferroni correction was applied.

Declaring a winner changes nothing about your test. Visitors continue to be served according to the configured traffic split, and the numbers keep updating, but the result is not expected to change. Rolling the winning variation out is a decision you make, in the A/B test configuration.