What it means
Two groups of customers average different spending. Is the gap a discovery or a coincidence?
The t-test is the century-old tool for answering. The logic compares signal to noise: the difference between means is weighed against the variation inside the groups, and big, quiet groups make small differences meaningful.
The NIST statistics handbook presents the forms: one-sample tests against a target, two-sample tests between groups, and paired tests when each subject is measured twice. The t distribution exists because of small samples: William Gosset, working at the Guinness brewery as Student, showed how uncertainty about the spread widens the test's tails when data is scarce.
The output is a p-value: the probability of seeing a gap this large if there were truly no difference, the number everyone quotes and half of everyone misreads. The assumptions carry the weight: roughly normal data, comparable spreads for the classic form, and independent observations, with violations quietly inflating the discovery rate.
The misuses are legendary: testing twenty questions and celebrating the one hit, peeking until significance appears, and confusing statistical significance with a difference that matters. For a non-finance reader, a t-test is the sober friend at the tasting: yes, this batch scored higher, but with this few sips and this much variation, would the same gap appear on a rerun?
Welch's variant is the practical default: it drops the equal-spread assumption and behaves better when the two groups vary by different amounts. Effect size travel with the test: Cohen's d translates the gap into spread units, answering the question the p-value cannot, whether the difference is worth money.
Multiple comparisons are the field's open wound: run enough tests and false hits pile up by design, so honest pipelines correct for the number of questions asked. The t-test's family is larger than its name: analysis of variance extends the same signal-to-noise logic to many groups at once, and the t is the two-group special case.
In practice
Real-world examples.
Example
A 2.8% versus 2.2% conversion gap on 4,000 visitors per arm returns a p-value of about 0.09, and the launch is postponed. The team treats the result as unproven rather than as a loss. It plans a larger test.
Example
The same gap at 40,000 visitors per arm clears significance, and the button ships. The extra data shrinks the standard error, so a gap of the same size is now far outside what chance would produce. The team also records the size of the improvement.
Example
A testing charter bans peeking and pre-sets sample sizes, the doctrine against self-deception. Analysts agree in advance what difference is worth detecting. Results are read only once the planned sample is complete.
Formula
Calculation
The t statistic equals the difference in means divided by the standard error of that difference; with roughly normal data it follows the t distribution with degrees of freedom set by sample size, and the p-value reads off the tail probability.
Worked example. A fictional shop compares the average order value of two customer groups, 10 orders each. Group A averages $52 and Group B averages $46, and each group has a standard deviation of $8.
- Standard error of the difference = sqrt(8^2 / 10 + 8^2 / 10) = sqrt(6.4 + 6.4) = sqrt(12.8) = $3.58.
- t = ($52 - $46) / $3.58 = 1.68.
- With about 18 degrees of freedom, the two-sided 5% critical value is 2.101, and 1.68 is below it, so the gap is not statistically significant.
- Effect size = $6 / $8 = 0.75 standard deviations, a moderate gap that the small samples cannot confirm.Case study
Seen in the real world.
This case study is fictional and illustrative. A made-up e-commerce team tests a green checkout button against the blue one on 4,000 visitors each and finds green converts 2.8% against blue's 2.2%. The product manager wants to ship by lunch; the analyst runs the two-sample test. The verdict is a p-value of about 0.09: a gap this size would appear about nine times in a hundred by chance alone, so lunch becomes a longer test, and the product manager learns the difference between a number and a result.
The rerun with 40,000 visitors per arm finds the same gap, now with a p-value far below 0.01, and the button ships with a confidence the first week could not have bought at any price. The team's testing charter, written after the episode, bakes in the doctrine: sample sizes decided before the test starts, no peeking at p-values mid-flight, and a minimum effect worth caring about set in advance, because a big enough sample can make a trivial difference significant. The analyst's poster above the experiment dashboard carries Gosset's own career as the moral: the test was invented by a brewer who could not afford large samples, which is exactly who still needs it. The green button is still there, joined by a graveyard of ideas that failed their t-tests quietly and cheaply.
Watch out
Common mistakes.
- Reading p as the chance the effect is real; it is the chance of data this extreme under no effect, a different and slipperier quantity.
- Ignoring assumptions; non-normal data, unequal spreads, and dependent observations all break the textbook form.
- Confusing significance with importance; large samples certify tiny effects, so report the size of the difference alongside the p-value.
Questions
People also ask.
What is a t-test?
A statistical test comparing means, weighing the difference between groups against the variation within them, built for small samples.
Who invented it?
William Sealy Gosset, a Guinness brewer publishing as Student in 1908, to handle quality decisions on small batches.
What does the p-value mean?
The probability of observing a difference at least this large if the true difference were zero; small values argue the gap is not noise.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%