What it means
Winsorizing started as a way to make the ordinary sample mean less sensitive to extreme values. You sort the data, then replace the smallest k values with the (k+1)st smallest value.
You do the same at the top, replacing the k largest with the (k+1)st largest, and then take the mean of the new set. SAS documentation describes trimmed and Winsorized means as estimators of the population mean that are relatively insensitive to outlying values.
The trimmed mean deletes the k smallest and k largest observations. The Winsorized mean keeps them in the sample but pulls them in.
A SAS statistician's article, which cites the Encyclopedia of Statistical Sciences, stresses two points. Winsorization is symmetric, so it changes the same number of values in each tail.
It is also based on counts, not on percentiles, because percentile cutoffs can change more values in one tail than the other when data repeat. The same sources note that for data from a symmetric population, the Winsorized mean is an unbiased estimate of the population mean.
If the data are skewed, the result can be biased. In finance, analysts often Winsorize returns, valuation ratios or earnings surprises before averaging or running a regression.
The choice of k is a judgment, so it should be stated and tested. A Winsorized mean is like asking the loudest and quietest people in a room to speak at the volume of their nearest neighbours.
Everyone still gets a voice, but nobody can drown out the rest.
In practice
Real-world examples.
Example
A fictional team records eight monthly expense growth rates in percent: 12, 14, 15, 15, 16, 17, 18 and 95. The 95 is a data entry error that should have been 19. The raw mean is 25.25, which describes none of the months.
Example
A fictional analyst Winsorizes the same eight values with k = 1. The 12 becomes 14 and the 95 becomes 18, giving 14, 14, 15, 15, 16, 17, 18, 18. The Winsorized mean is 127 / 8 = 15.875, close to the median of 15.5.
Example
A fictional researcher studies price-to-earnings ratios for 500 stocks. A handful of firms with near-zero earnings show ratios above 1,000. She Winsorizes the top and bottom 5 values before averaging, and reports the choice of k = 5 in the methods note.
Formula
Calculation
Sort the data as y(1) <= y(2) <= ... <= y(n). For a k-times Winsorized mean, replace y(1) to y(k) with y(k+1), and replace y(n-k+1) to y(n) with y(n-k).
Mean = [ (k+1) x y(k+1) + sum of y(k+2) to y(n-k-1) + (k+1) x y(n-k) ] / n.
Worked example with assumed figures: n = 8 and k = 1. The sorted values are 12, 14, 15, 15, 16, 17, 18, 95. The 12 is replaced by 14 and the 95 by 18, giving 14, 14, 15, 15, 16, 17, 18, 18, so the sum is 127.
The Winsorized mean is 127 / 8 = 15.875.
For comparison, the raw mean is 202 / 8 = 25.25. The trimmed mean with k = 1 drops 12 and 95 and averages the remaining six values, 95 / 6 = 15.83. The median is 15.5.Case study
Seen in the real world.
This case study is fictional and illustrative. A credit risk analyst tracks monthly loss rates on eight loan pools. Seven pools sit between 12 and 18 basis points, and one shows 95 because a servicing file was loaded twice. The raw mean of 25.25 would make the portfolio look much worse than it is.
The analyst Winsorizes with k = 1 and gets 15.875, then compares it with the median of 15.5 and a trimmed mean of 15.83. The three outlier-resistant measures agree within half a basis point. She then investigates the 95 and finds the duplicate load. After the fix, the true value is 19, and the raw mean becomes 15.75.
The note records that Winsorizing was a screening step, not a replacement for fixing the data. Outlier-resistant averages help spot a problem, but they do not decide whether an extreme value is real. A genuine extreme loss belongs in a risk report.
Watch out
Common mistakes.
- Winsorizing at percentiles without checking how many values change in each tail, which can break the symmetry the method relies on.
- Treating Winsorizing as data cleaning, when it hides errors and real extremes alike. Investigate the extreme values first.
- Using it on skewed data and calling the result unbiased, when the unbiasedness claim is for symmetric populations.
Questions
People also ask.
What is a Winsorized mean?
It is the average of a data set after the k smallest and k largest values are replaced by their nearest remaining neighbours. No observations are removed.
How is it different from a trimmed mean?
A trimmed mean deletes the extreme observations, so the sample shrinks. A Winsorized mean replaces them, so the sample size stays the same.
How do I pick k?
There is no universal rule. Analysts often use a small share of the sample at each end, state it openly and check that results do not change much for nearby values.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%