What it means
Penn State's STAT 200 course defines a Type II error as failing to reject the null hypothesis when the null hypothesis is really false, denoted by beta, the probability of a Type II error. The matching Type I error is rejecting the null hypothesis when it is really true, with probability alpha.
The course sets this out as a two-by-two table in which rejecting a true null and failing to reject a false null are the errors, and the other two cells are correct decisions. The course gives a courtroom example in which the null hypothesis is not guilty, so a Type II error means the man did kill his wife but was found not guilty and was not punished.
A second example involves asparagus: culinary students test whether more than half of patrons prefer a frying method, and a Type II error means they conclude the new method is not superior when it really is. A study of retail customers illustrates the same thing, since the sample proportion was 0.48 while the true proportion was 0.53, so not rejecting was a Type II error.
Penn State defines power as the probability of correctly rejecting a false null hypothesis, and power equals 1 minus beta. So when power rises, the Type II error probability falls.
The course lists ways to increase power, including a larger sample size, a smaller standard error, a bigger difference between the sample statistic and the hypothesised value, and a larger alpha. A directional test also has more power than a two-tailed test.
There is a trade-off, because with a fixed sample size, decreasing alpha increases beta. The course says that to lower both error probabilities, increase the sample size.
The course also warns that failing to reject the null does not mean accepting it. Either the null is really true or the sample was too small to reject it, meaning power was too low.
In practice
Real-world examples.
Example
A fictional bank tests whether a new fraud rule raises the detection rate. The rule truly works, but the sample is small and the test does not reject the null. The bank wrongly concludes the rule adds nothing.
Example
A fictional retailer tests whether a loyalty program raises average basket size. The true effect is positive, but the test uses a strict alpha of 0.01 on a small sample. The test fails to reject, and a useful program is dropped.
Example
A fictional lender tests whether a new credit score cuts defaults. The data show a small drop that is not statistically significant. The lender should say it lacked enough evidence, not that the score does nothing.
Formula
Calculation
Power = 1 - beta.
Example with assumed figures: test whether a mean exceeds 100 with a standard deviation of 15, a sample of 36 and alpha of 0.05 (one-sided). The standard error is 15 / sqrt(36) = 2.5, so the cutoff is 100 + 1.645 x 2.5 = 104.11.
If the true mean is 106, beta = P(sample mean < 104.11) = about 0.225, so power is about 0.775.
Raising the sample to 64 cuts beta to about 0.060. Cutting alpha to 0.01 at n = 36 raises beta to about 0.471. These were computed in Python.Case study
Seen in the real world.
This case study is fictional and illustrative. An insurer pilots a new claims triage model. It wants to know if the model reduces average handling time below the current 100 minutes. The analyst sets alpha at 0.05 and uses a pilot of 36 claims, with a standard deviation of 15 minutes. The test is set up to detect a true shift of 6 minutes.
Using the figures above, the chance of missing a real 6 minute improvement is about 22.5 percent. That is the Type II error probability, and the power is about 77.5 percent. The analyst reports this beside the result. The pilot result is not significant. Rather than saying the model fails, the team notes the low power and extends the pilot to 64 claims.
At that size beta falls to about 6 percent. The lesson is to plan sample size for power before the test starts. A non-significant result from an underpowered test says little.
Watch out
Common mistakes.
- Saying a non-significant result proves the null hypothesis, when Penn State says failing to reject is not accepting.
- Ignoring the alpha and beta trade-off, since lowering alpha at a fixed sample size raises beta.
- Planning sample size for alpha only, when power needs to be planned in advance.
Questions
People also ask.
What is a Type II error?
It is failing to reject a null hypothesis that is really false. The real effect is missed.
How is it related to power?
Power is 1 minus beta, where beta is the Type II error probability. Higher power means a lower chance of a Type II error.
How can I reduce the chance of a Type II error?
Increase the sample size, reduce the standard error, or accept a larger alpha. Using a directional test also adds power when the question is one-sided.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%