What it means
Imagine you test a new checkout page and conversion rises from 10% to 11%. The question is whether the page worked or whether you were simply lucky with the particular customers who visited.
The p-value helps answer this by asking how often a gap that large would appear by chance alone. The starting assumption is called the null hypothesis, which says there is no real effect or difference.
If the p-value is small, the data would be surprising under that assumption, and you may decide to reject it. A common cut-off is 0.05, meaning a 5% chance, but the right threshold depends on the cost of being wrong.
For business users the meaning is easily misread. A p-value does not tell you the probability that your idea is right, and it does not measure the size or importance of an effect.
A tiny effect can have a very small p-value if the sample is huge, and a big effect can have a large p-value if the sample is small. Finance and analytics teams use p-values to test things like whether a trading strategy beats the market, whether a price change altered demand, or whether a fraud model finds more cases than random checking.
They are also seen in regression output, where each variable gets its own p-value. They should be read together with the size of the effect and the cost of acting on it.
Good practice is to decide the threshold before looking at the data, report the effect size as well, and avoid running many tests until one looks significant. Doing so inflates false positives, because with enough tests some will pass by luck.
In practice
Real-world examples.
Example
A marketing team compares two email subject lines. One has a 3.2% click rate and the other 3.0%, and the test gives a p-value of 0.40. The team concludes the difference is probably noise and does not change its approach.
Example
A fund manager claims to beat the index by 1.5% a year. An analyst tests ten years of returns and finds a p-value of 0.30. The track record is not strong enough to rule out luck.
Example
A retailer raises its prices by 4% in a pilot region and tests whether sales volume changed. The p-value is 0.01, so the fall in volume is very unlikely to be chance. The team decides whether the higher price makes up for it.
Formula
Calculation
Z = (sample mean - hypothesised mean) / (standard deviation / square root of sample size)
Two-sided p-value = probability of a z-score at least this far from zero in either direction
An online shop's average order has historically been $50. After a redesign, a sample of 100 orders averages $52, and the standard deviation of order values is $10. Standard error = 10 / square root of 100 = 10 / 10 = 1. z = (52 - 50) / 1 = 2.0. The two-sided p-value for z = 2.0 is about 0.0455, or 4.55%.
Reading the result: if the redesign had no real effect, a sample average at least $2 away from $50 would turn up in only about 4.55% of samples. That is below the usual 0.05 cut-off, so the result is called statistically significant, but it does not prove the redesign caused the rise or that the extra $2 per order is worth the cost.Case study
Seen in the real world.
Cobalt Pay is an illustrative, fictional payments app that tested a new sign-up screen on 400 users while 400 others saw the old one. The new screen converted 46% against 40% for the old one, and the product manager wanted to launch immediately.
The analyst calculated a p-value of 0.09, above the team's pre-agreed threshold of 0.05. She explained that a gap of six points could arise by chance about nine times in a hundred even if the screens were equal.
The company extended the test to 2,000 users on each side, and the gap narrowed to 2 points with a p-value of 0.20. The illustrative lesson is that early excitement from a small sample often fades when more data arrives.
Watch out
Common mistakes.
- Reading a p-value of 0.03 as a 97% chance that the idea works, when it is not a probability about the idea at all.
- Treating a small p-value as proof of a large effect, when a huge sample can make a trivial difference look significant.
- Running many tests and reporting only the ones with a p-value under 0.05.
Questions
People also ask.
What is a good p-value?
There is no universal rule, though 0.05 is a common cut-off and some fields use stricter limits such as 0.01.
Does a high p-value prove there is no effect?
No, it only means the data do not provide strong evidence of one, and a larger sample might show something.
What does statistically significant mean?
It means the p-value fell below the chosen threshold, not that the result is important for the business.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%