What it means
Many familiar tests, such as the t-test, assume that the data are roughly normally distributed, which means they cluster symmetrically around an average. Business data often break this assumption, since figures such as invoice values, delivery times and customer ratings tend to be skewed by a few extreme results.
The Wilcoxon test avoids the problem by using ranks, which means the values are lined up from smallest to largest and numbered. A single extreme value then counts only as the highest rank, not as a huge number that pulls the average.
There are two main versions. The signed-rank test is for paired data, where the same item is measured twice, such as a branch before and after a process change.
The rank-sum test, also known as the Mann-Whitney test, is for two independent groups, such as customers in two regions. For the signed-rank test, you calculate the difference for each pair, ignore the signs and rank the sizes, then add up the ranks of the positive differences and of the negative differences separately.
The smaller total is the test statistic, and a very small value suggests that the changes mostly go in one direction. The result is compared with a table or software to produce a p-value, which is the chance of seeing such a result if there were no real difference.
A low p-value, commonly below 0.05, is taken as evidence of a genuine difference, but with small samples the test has limited power. Managers should treat the output as one input to a decision.
A significant result says that a difference is unlikely to be chance, and it does not say how large or valuable the difference is.
In practice
Real-world examples.
Example
A retailer measures the weekly sales of 12 stores before and after a new layout and finds that the changes are uneven, with a few very large increases. It uses the signed-rank test instead of a t-test because the differences are skewed. The test shows whether most stores improved.
Example
A call centre compares customer satisfaction scores from 1 to 5 between two teams. The scores are ranked categories rather than true measurements, so an average is not reliable. The rank-sum test checks whether one team tends to score higher.
Example
An accounts payable manager compares invoice processing times for two suppliers with only eight invoices each. The sample is too small to trust the normal curve assumption. The test gives an evidence-based answer without relying on that assumption.
Formula
Calculation
W+ = Sum of ranks of positive differences; W- = Sum of ranks of negative differences; Test statistic W = smaller of W+ and W-
A finance team introduces a new approval workflow at six branches and records the change in days to approve an invoice, where a positive number means faster. The differences are +3, +5, -1, +4, +2 and +6 days. Ranking the sizes 1, 2, 3, 4, 5, 6 from smallest to largest gives rank 1 to the -1, rank 2 to the +2, rank 3 to the +3, rank 4 to the +4, rank 5 to the +5 and rank 6 to the +6. W+ = 3 + 5 + 4 + 2 + 6 = 20 and W- = 1, so W = 1. With six pairs there are 64 equally likely sign patterns, and only 2 of them give W- of 1 or less while 2 more give W+ of 1 or less, so the two-sided p-value = 4 / 64 = 0.0625. That is just above 0.05, so the improvement is suggestive but not conclusive.Case study
Seen in the real world.
Kestrel Logistics is an illustrative, fictional courier firm that tried a new route-planning tool at ten depots. It recorded the average delivery time in minutes before and after the change for each depot.
Nine depots improved and one worsened, but the sizes of the changes varied widely, and one depot improved by far more than the others. The analyst chose the signed-rank test because that one depot would have distorted an average, and the result gave W- = 3 with a p-value below 0.05.
In the illustrative outcome, the company concluded that the tool was likely to help and rolled it out to the remaining depots. The analyst also reported the median saving of 6 minutes per delivery, so the board understood the size of the benefit as well as the statistical result.
Watch out
Common mistakes.
- Using the rank-sum version on paired data, or the signed-rank version on independent groups, which gives misleading answers.
- Treating a p-value below 0.05 as proof that a difference is important, when it only shows it is unlikely to be chance.
- Ignoring ties and zero differences, which need special handling in the calculation.
Questions
People also ask.
When should I use the Wilcoxon test instead of a t-test?
When the sample is small, the data are skewed or the values are ranks, so the normal curve assumption is doubtful.
What is the difference between the Wilcoxon and Mann-Whitney tests?
The signed-rank Wilcoxon test is for paired data, while the rank-sum version, also called the Mann-Whitney test, is for two independent groups.
Does the test tell me how big the difference is?
No, it tells you whether a difference is likely to be real, so you should also report a measure such as the median change.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%Related
