What it means
Almost every business number moves around from month to month for random reasons. If conversion rates rise from 3.0% to 3.2% after a website change, you need a way to tell whether the change worked or the rise was luck.
Statistical significance gives a disciplined answer to that question. The method starts with a null hypothesis, which is the assumption that nothing has changed.
The analyst then calculates a p-value, the probability of seeing a result at least as extreme as the one observed if the null hypothesis were true. If that probability falls below a chosen cut-off, often 5%, the result is called statistically significant.
The cut-off, known as the significance level, is a choice rather than a law of nature. A lower cut-off such as 1% demands stronger evidence and suits high-stakes decisions.
A higher one is quicker to trigger action but increases the chance of false alarms. Significance is heavily affected by sample size.
With millions of observations, a trivially small difference can be statistically significant, while with only a handful of data points even a big difference may not be. This is why managers also ask about effect size, meaning how large the difference is in business terms.
A further nuance is that significant does not mean certain. At a 5% cut-off, roughly one in twenty tests of a change that truly does nothing will still look significant by accident.
Running many tests and reporting only the winners is a well-known way to fool yourself. For a non-specialist, the practical habit is to ask three questions whenever someone claims a significant result.
How big is the difference in dollars, how many observations sit behind it, and was the test planned before the data was seen? If those answers are weak, the claim deserves more scepticism no matter how impressive the p-value looks.
In practice
Real-world examples.
Example
A marketing manager tests two email subject lines on 10,000 customers each. One gets a 2.5% click rate and the other 2.9%, and the analyst reports the gap is statistically significant. The team adopts the better line for the next campaign.
Example
A fund analyst finds that a trading strategy beat the market by 0.4% a year over three years. A significance test shows the outperformance could easily be luck given the volatility of returns. The investment committee declines to allocate capital.
Example
A hospital group trials a new billing process in one clinic and cuts average claim processing time from 12 days to 11 days. With only 25 claims in the sample, the result is not statistically significant. Management extends the trial to more clinics before deciding.
Formula
Calculation
Z = (sample mean - hypothesised mean) / (standard deviation / square root of sample size)
Suppose a retailer's average order value has long been $80 with a standard deviation of $20. After a redesign, a sample of 100 orders averages $84. The standard error is 20 / 10 = $2, so z = (84 - 80) / 2 = 2.0. A z-score of 2.0 corresponds to a two-sided p-value of about 4.6%, which is below 5%, so the increase is statistically significant at the 5% level.Case study
Seen in the real world.
Peakline Fitness is an illustrative, fictional chain that tried a new pricing page on its website. In the first week, sign-ups rose 8% compared with the old page, and the marketing lead wanted to roll it out immediately.
The analyst noted that the week included only 150 sign-ups on each version, and a significance test gave a p-value of 31%, well above the 5% cut-off. The company kept the test running for six weeks until it had several thousand visitors on each page.
At that point the increase was only 2%, but it was statistically significant and worth about $90,000 a year. The illustrative lesson is that early excitement often comes from small samples, and patience protects against costly mistakes.
Watch out
Common mistakes.
- Equating statistically significant with important, when a tiny difference can be significant in a large sample and still be commercially worthless.
- Stopping a test the moment the p-value dips below 5%, which inflates the false positive rate.
- Reading a p-value of 4% as a 96% chance the result is real, when it is the chance of such data if nothing had changed.
Questions
People also ask.
What is the usual significance level?
Many fields use 5%, but 1% or 10% are also used depending on the cost of being wrong.
Can a result be real but not statistically significant?
Yes, if the sample is too small to detect it, which is why more data is sometimes needed.
What is a confidence interval and how does it relate?
It is a range of plausible values for the true effect, and a 95% interval that excludes zero corresponds to significance at the 5% level.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%