What it means
A sample average can differ substantially from its expected value when only a few observations are available. The law describes convergence as the number of observations grows, not a promise that the next observation will correct the previous error.
The path of an average can move toward or away from the target along the way. Independence means knowing one observation does not provide information about the others in the model.
Identically distributed means each observation follows the same probability distribution. These assumptions describe an idealised sampling or repeated-experiment process rather than every collection of business data.
The University of Maryland notes explain the weak law in terms of probability. For any chosen positive tolerance, the probability that the average differs from the expected value by more than that tolerance approaches zero as the observation count increases.
This is a limiting statement, not a fixed sample-size rule. Increasing volume does not cure selection bias: a survey that excludes dissatisfied customers can produce a stable but misleading average, because more observations from the wrong group reinforce the wrong measurement rather than establish the average for the population the manager intended to study.
Dependence and changing conditions can also matter, since orders from one customer during one promotion may share a common influence, while production measurements before and after a machine change may follow different distributions. Pooling them requires more care than simply counting the rows.
The law differs from the gambler's fallacy, because several heads in independent coin tosses do not make a tail more likely on the next toss. Long-run convergence of the proportion does not create a balancing force that changes the probability of the next independent event.
It also differs from the central limit theorem, which concerns the distribution of appropriately scaled deviations under its own conditions, while the law of large numbers concerns convergence of the average. Neither result says that every original dataset becomes bell-shaped merely because it grows.
Business commentary sometimes uses the phrase for slower percentage growth at larger companies, but that colloquial meaning is not the probability theorem.
In practice
Real-world examples.
Example
A fictional analyst simulates independent coin tosses with a fixed heads probability. The proportion can fluctuate at each step even though the model's long-run convergence result applies. A temporary movement away from the target does not contradict the theorem.
Example
A retailer gathers thousands of responses only from customers who completed a purchase. The analyst explains that a large response count cannot by itself reveal the views of people who abandoned the checkout.
Example
A factory combines measurements from several machines after one was recalibrated. Before interpreting the overall average, the team examines whether the observations still represent the same process rather than assuming the row count resolves the difference.
Formula
Calculation
Sample mean = sum of observed values / number of observations. For a simple indicator, record 1 for an event and 0 otherwise; the mean then equals the observed event proportion.
In a fictional set of 100 independent simulated trials, 54 events give a proportion of 0.54. After another 900 trials, the combined total is 510 events out of 1,000, giving 0.51. This example illustrates possible behaviour around a modelled probability of 0.50, not a guarantee that the next trial or every larger sample must move closer.Case study
Seen in the real world.
In this fictional case, Harbour Analytics reports an average delivery delay from a large shipment file. Management assumes the record count makes the figure automatically reliable and asks to use it for every customer segment. The analyst checks the sampling process and finds that several major customers use a different delivery service, while some late shipments were omitted from the file.
The team separates the populations and repairs the missing-record problem before interpreting the averages. The report separates quantity, quality, and assumptions. The case shows why a convergence theorem can support analysis without replacing the need to define what was observed and which population the estimate describes.
Watch out
Common mistakes.
- Assuming every new observation must move the average closer to the expected value.
- Treating a large biased or dependent dataset as automatically representative of the intended population.
- Using the theorem to claim that previous independent outcomes change the probability of the next event.
Questions
People also ask.
Does the law guarantee a correct finite-sample estimate?
No. It describes convergence under specified assumptions as the observation count grows. A particular finite sample can still differ from the expected value.
Do repeated heads make a tail due next?
No. Under independence, earlier results do not change the next toss's probability. Long-run averaging is different from forced short-run balancing.
Can more data fix a biased sampling method?
Not by itself. The observations must support the intended population and model. More data from an excluded or selected group can preserve the same bias.
From the founder's library

Take it further with the book.
Build your financial confidence beyond this definition. Shihan's full-length guide, Accounting Fundamentals, takes the same plain-English approach and turns it into a complete, practical playbook for non-finance managers, business owners and students - with chapter-end quiz answers and presentation slides included.
25% off with code MMHQ25, applied at checkout. Priced in USD - checkout may show the equivalent in your local currency.
View the book and save 25%