Back to Glossary

Bonferroni Test

A Bonferroni test applies a stricter threshold to each of several statistical comparisons so that the probability of at least one false positive across a defined family stays within a chosen limit. The simplest version divides the familywise significance level by the number of tests.

It is useful when analysts test many business hypotheses, but its conservatism can also hide genuine effects.

From the Money Master HQ dictionary, founded by Shihan Sheriff (FCMA, VP of Finance at Nomod, CFO at Esanjo Ventures). How these definitions are written.

What it means

A single test at a 5% significance threshold has a controlled chance of rejecting a true null under its assumptions. If an analyst looks at many unrelated metrics and reports only a significant one, the chance of at least one false alarm rises.

The relevant question is how many tests belonged to the decision, not how many attractive results were published. Bonferroni sets the per-test threshold at alpha divided by m, where m is the number of comparisons in the family.

For ten tests and a desired familywise alpha of 0.05, each test must meet 0.005. A p-value of 0.03 would not pass that corrected threshold, even though it is below 0.05.

The family should be defined before interpreting results. If a retailer checks ten store promotions as one analysis, excluding nine from the correction after seeing the outcomes defeats its purpose.

Unrelated research programmes need not automatically be lumped together, but the rationale for grouping must be defensible. The method controls the chance of one or more false rejections using an inequality that does not require test independence, which is useful when metrics are correlated.

It does not repair bad sampling, biased measurement, a misspecified model, or repeated unreported attempts to adjust the data. A smaller threshold also makes detection harder, so with many tests or small samples, meaningful effects may not reach the corrected criterion.

This is a trade-off between false positives and false negatives, not evidence that a result with p = 0.006 has no business value. Other methods, including Holm's step-down method and false-discovery-rate controls, answer related but different error-control questions.

Choice depends on how costly a single false claim would be and how many true findings the team hopes to discover, and it should not be made after seeing which method makes a favourite conclusion significant. The correction does not multiply or divide the effect itself: a $20 lift per customer remains an estimated $20 lift, and only the evidence threshold changes.

Report effect size, uncertainty interval, sample size, and design alongside the corrected decision rather than presenting a binary label alone.

In practice

Real-world examples.

1

Example

A firm tests ten advertising messages for conversion lift and wants no more than 5% familywise false-positive risk under its test assumptions. The per-message Bonferroni threshold is 0.05 / 10 = 0.005. One message with p = 0.004 passes; another with p = 0.02 does not, though both merit effect-size review.

2

Example

An analyst tests 20 trading signals, reports only the smallest p-value, and calls it a 5% discovery. The research lead requests the entire test set and applies a 0.05 / 20 = 0.0025 threshold for this family. The result's p-value of 0.01 is not convincing under that chosen control.

3

Example

A hospital service team evaluates five separate safety outcomes after a process change. A moderate but important effect misses the stricter threshold. Managers do not call the effect zero; they report its estimate and uncertainty, consider a larger sample, and avoid overstating evidence.

Formula

Calculation

Bonferroni per-test threshold = familywise alpha / number of tests. Worked example: a team tests ten advertising messages with a familywise alpha of 0.05, so the threshold is 0.05 / 10 = 0.005. Message A has a raw p-value of 0.004, which is below 0.005, so it passes. Equivalently, its adjusted p-value is 0.004 x 10 = 0.04, which is below 0.05. Message B has a raw p-value of 0.02, which exceeds 0.005; its adjusted p-value is 0.02 x 10 = 0.20, so it does not pass. Adjusted p-values are capped at one, and the calculation must use the actual prespecified family size.

Case study

Seen in the real world.

Fictional example: Solara Retail tested ten checkout messages across comparable customer groups. Analyst Noor noticed one result with p = 0.004 and another with p = 0.02. A presentation draft labelled both statistically significant at 5%, although the decision was based on the whole set. Noor listed all ten prespecified messages and calculated the Bonferroni threshold of 0.005; only the first passed.

She also reviewed conversion differences, sample sizes, and whether customer allocation had been sound. The second result still informed a future test but was not announced as a confirmed lift. The team rolled out the first message in a limited monitored trial and saved the complete experiment report. It did not quietly redefine the family after seeing which promotions performed best.

Noor added a short note explaining why the correction mattered. If none of the ten messages had any real effect and the tests were independent, the chance of at least one result below 0.05 would be 1 - 0.95^10, or about 40%. That is why a lone small p-value from a large set deserved caution.

Watch out

Common mistakes.

  • Applying 0.05 separately to many selected tests and reporting only the winners.
  • Treating a result that misses the adjusted threshold as proof the underlying effect is exactly zero.
  • Believing multiple-testing correction fixes sampling bias, a bad model, or hidden repeated analyses.

Questions

People also ask.

Does Bonferroni require independent tests?

No. The familywise bound works without independence, although it can be conservative when tests overlap.

How many tests belong in m?

Use the defined family of comparisons relevant to the claim, including results that were not significant.

Is it always the best correction?

No. Its strictness can reduce power. Choose an error-control method based on the decision and prespecified analysis plan.

Was this explanation helpful?

From the founder's library

Accounting Fundamentals: A Non-Finance Manager's Guide to Finance and Accounting, by Shihan Sheriff

Take it further with the book.

Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.

US$2.24US$2.99

25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.

View the book and save 25%
Last updated · October 8, 2026
Browse all terms →

Disclaimer

The information provided in this finance dictionary is for educational and informational purposes only. It should not be construed as financial, investment, legal, or tax advice. Always consult with a qualified professional before making any financial decisions. Money Master HQ makes no representations or warranties about the accuracy, completeness, or suitability of this information. Use of this content is at your own risk.