Back to Glossary

Hypothesis Testing

Hypothesis testing is a statistical process for assessing whether data provide enough evidence against a specified null hypothesis. It compares observed results with what would be expected under that hypothesis, using a chosen decision rule and assumptions about how the data were generated.

From the Money Master HQ dictionary, founded by Shihan Sheriff (FCMA, VP of Finance at Nomod, CFO at Esanjo Ventures). How these definitions are written.

What it means

The null hypothesis states the reference claim being tested, such as no difference between two population averages, while the alternative describes the difference or direction the analysis is designed to detect. The test does not begin by proving the alternative true.

It asks whether the observed data would be sufficiently unusual if the null and the model assumptions were correct. A test statistic summarises the relevant evidence, and its reference distribution determines how extreme the observed result is under the null hypothesis.

A p-value measures the probability, under the specified null model, of obtaining a result at least as extreme as the one observed. It is not the probability that the null hypothesis is true.

The significance level is chosen before looking at the result. It sets the decision threshold and relates to the risk of a false rejection when the null is true under the model.

A Type I error rejects a true null, while a Type II error fails to reject a false one, and sample size, effect size, variation and the test design affect the balance between these risks. Statistical significance is different from commercial importance, because a large sample can make a tiny improvement statistically detectable even when it has little value after implementation costs.

Failure to reject the null does not prove that there is no effect. The study may be too small, noisy or poorly designed to detect a meaningful difference.

NIST's statistical guidance distinguishes p-values, critical values and decision thresholds. The result should be interpreted alongside estimates and uncertainty intervals rather than reduced to one pass-or-fail label.

Managers should also check data quality and design, since a precise test cannot repair biased sampling, changed measurement or comparisons selected only after examining many possible results.

In practice

Real-world examples.

1

Example

A retailer tests whether a new checkout process changes average waiting time. It sets the outcome and test approach before reviewing the results, rather than searching afterwards for whichever metric improved.

2

Example

A very large customer study finds a statistically significant satisfaction increase of 0.1 points. Management compares that small effect with training costs before rolling out the change.

3

Example

A pilot shows no statistically significant sales lift, but the sample is small. The team reports uncertainty rather than claiming the new offer definitely has no effect.

Formula

Calculation

For a simple one-sample mean test, a t statistic is sample mean minus hypothesised mean, divided by the sample standard deviation over the square root of sample size. The appropriate assumptions and reference distribution must hold. Suppose 25 observations have mean 104, sample standard deviation 10 and a null mean of 100. The standard error is 10 / 5 = 2, giving a t statistic of (104 - 100) / 2 = 2. With 24 degrees of freedom, the two-sided 5% critical value is about 2.064, so a t statistic of 2 would not lead to rejection at that level, while a one-sided test with a critical value of about 1.711 would. A 95% confidence interval for the difference is 4 plus or minus 2.064 x 2, or roughly -0.13 to 8.13, which includes zero. The statistic alone is not a complete conclusion, and the estimated difference of 4 units should also be assessed for practical importance.

Case study

Seen in the real world.

The following is an illustrative and fictional case. Clear Path Services tested a new customer reminder system in two comparable branches. The first report focused on a p-value below the agreed threshold and recommended immediate expansion. Finance asked for the size of the improvement, the implementation cost and whether the branches had comparable customer mixes. The review found that the effect was modest and that one branch had changed staffing during the pilot.

The team redesigned the test to separate the reminder effect from the operational change. The second study used a clearer comparison and reported both the estimated improvement and its uncertainty. Management rolled out the system only where expected savings justified the cost. The statistical test helped organise the evidence, but it did not replace the business decision. Design quality, effect size and implementation economics determined what the company did next.

Watch out

Common mistakes.

  • Reading a p-value as the probability the null is true. It is calculated under a specified null model.
  • Equating statistical significance with business value. Effect size and cost determine whether a change is useful.
  • Treating a non-significant result as proof of no effect. Limited data or high variation may conceal a meaningful difference.

Questions

People also ask.

What is the null hypothesis?

It is the reference claim the test evaluates, such as no difference or a specified population mean.

Should the threshold be chosen after seeing the data?

No. Choosing it afterwards can distort the error control and make a favourable result look more convincing than it is.

What else should accompany the test?

Report the design, assumptions, effect estimate and uncertainty, then explain the practical consequences rather than relying on the p-value alone.

Was this explanation helpful?

From the founder's library

Accounting Fundamentals: A Non-Finance Manager's Guide to Finance and Accounting, by Shihan Sheriff

Take it further with the book.

Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.

US$2.24US$2.99

25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.

View the book and save 25%
Last updated · October 8, 2026
Browse all terms →

Disclaimer

The information provided in this finance dictionary is for educational and informational purposes only. It should not be construed as financial, investment, legal, or tax advice. Always consult with a qualified professional before making any financial decisions. Money Master HQ makes no representations or warranties about the accuracy, completeness, or suitability of this information. Use of this content is at your own risk.