The first lesson left a debt. It claimed that measured frequencies stabilise, could not justify it, and took the axioms as a starting point instead. Everything since then has been built on those axioms, and they now suffice to pay the debt back with interest: not only do averages settle, but the residual error has a shape that can be computed, and the shape does not depend on what was being averaged.
What has to be proved
Take independent repetitions of the same experiment, each with mean and finite variance , and let
be the average of the first . The claim to be proved is that closes in on . Applied to indicators, where is when an event occurs and otherwise, is the observed frequency of that event and is its probability, so proving the general claim proves the specific one about coin tosses.
Two ingredients are already in hand. By linearity, for every , so the average is centred correctly from the very first trial. By the additivity of variance for independent variables and the scaling rule,
so the spread of the average shrinks towards zero. Centred correctly and shrinking: that is nearly the whole proof, and Chebyshev supplies the rest.
The weak law
Chebyshev's inequality, applied to with its mean and standard deviation , says that for any fixed tolerance ,
obtained by setting in the inequality. The right-hand side is a constant divided by , so it tends to zero. That is the weak law of large numbers: for every tolerance, however tight, the probability that the average misses the true mean by more than that tolerance can be driven as low as you like by taking enough trials.
The proof is three lines and it uses only results already established, which is worth pausing on. Bernoulli's original proof in Ars Conjectandi, published in 1713, ran to twenty pages and he worked on it for twenty years. Chebyshev's inequality of 1867 collapsed it, and the collapse is what a good abstraction buys.
Note what the theorem does not claim. It does not say eventually stays near ; for each separately it bounds a probability, and a rare excursion at some later is not excluded. Strengthening it to "with probability one, the sequence of averages converges to and stays there" is the strong law of large numbers, proved for coin tossing by Émile Borel in 1909 and in general by Kolmogorov in 1930. The strong law needs no variance at all, only a finite mean. Its proof is genuinely harder and is not attempted here; the statement is what matters, and it is what licenses the frequency reading of probability from the first lesson.
Example. A measurement has standard deviation unit. Using Chebyshev, how many repeats guarantee that the average is within units of the true value with probability at least ?
The bound is , and setting this to gives repeats. This is a guarantee for any distribution whatsoever, which is why it is so demanding.
Now you. With the same and , what does Chebyshev give as the bound on the failure probability after repeats?
Answer
, so at least a percent chance of landing within of the true mean.
What the law does not promise
The law of large numbers is the single most misquoted result in mathematics, and the misquotation has a name: the gambler's fallacy, the belief that a run of one outcome makes the other "due".
The trials are independent. A coin that has landed tails ten times running lands heads next with probability exactly , because the coin has no memory and nothing in the mathematics says otherwise. So how does the average recover?
By dilution, not by compensation. After ten tails you are ten heads behind. Toss a thousand more times and the expected number of heads among them is , leaving you at heads out of , a fraction of . Toss ten thousand more and it is out of , or . The deficit of ten never goes away; it is simply divided by an ever larger denominator. This is exactly the pattern in Kerrich's table from the first lesson, where the absolute surplus of heads grew while the fraction converged, and the two facts are now both consequences of with .
The second misuse is applying the law to a small number of trials. Nothing in it says anything about ten tosses, or a hundred, or an evening at a casino. The bound weakens by a factor of , so it says nothing useful until is large relative to .
The shape of what is left over
The weak law says the error shrinks. That leaves an obvious question: shrinks how, and what does it look like on the way?
The scale is already known. The standard deviation of the error is , so multiplying the error by produces something whose spread stays fixed at as grows. Equivalently, standardise the sum:
which has mean and variance for every . The central limit theorem says that as grows, the distribution of converges to the standard normal, whatever the distribution of the individual was, provided only that its variance is finite.
Read that again, because it is a strange claim. The ingredients can be dice, coin tosses, exponential waiting times, incomes, anything at all with a finite variance, symmetric or skewed, discrete or continuous. Sum enough of them, standardise, and the answer is the same curve to whatever accuracy you like. The individual distribution is forgotten entirely except through its mean and variance. Laplace stated a version in 1810, Lyapunov proved it under general conditions in 1901, and the name is Pólya's, from 1920.
This explains the observation the previous lesson could not. Measurement errors are normal because each is a sum of many small independent disturbances. Heights are roughly normal because many small genetic and environmental contributions add. And de Moivre's binomial approximation is the special case where the are Bernoulli trials, since a binomial variable is literally a sum of of them.
Example. Ten fair dice are rolled and totalled. The single-die mean is and variance , so the total has mean and standard deviation . Estimate using the central limit theorem with a continuity correction, and compare with the exact value.
, so the estimate is . The exact answer, computed by convolving the six-sided distribution ten times, is . Ten dice is already enough for agreement to three decimal places.
Now you. Using the same figures, estimate the probability that the total of ten dice is at least .
Answer
, so the estimate is . The exact value is , so the approximation is a little over one percent low.
What a poll is actually claiming
The most familiar application is the reported margin of error on an opinion poll, and it is now derivable in full.
Poll people independently, each supporting a proposition with unknown probability . The count is binomial, and the observed proportion has mean and standard deviation . That quantity is largest at , where it equals , so using the worst case costs little and needs no knowledge of . By the central limit theorem is approximately normal, and percent of a normal distribution lies within standard deviations of its mean. Hence with probability about ,
For this is , the familiar "plus or minus three points". Quadrupling the sample to halves it to points, which is the square root law again and is why polls are not larger: the fourth thousand respondents buy far less than the first thousand.
Two honest caveats. The figure assumes a genuine random sample, and in practice non-response and coverage errors dwarf the sampling error this formula describes, so a poll's real uncertainty is wider than its stated margin. And the margin is a statement about the procedure, not about this particular poll: one poll in twenty is expected to fall outside its own margin, which is worth remembering when a single surprising result appears.
Example. A poll of people is conducted. What is the worst-case percent margin of error?
, about percentage points.
Now you. How many people must be polled for a worst-case margin of percentage point?
Answer
Set , so and people.
Where the limit laws fail
Both theorems have conditions, and each condition fails somewhere real.
Independence is the first. Correlated observations do not average out at the rate , and if the correlation does not decay with distance they may not average out at all. Sampling a thousand people from one town is not sampling a thousand people from the country, and treating it as such gives a margin of error that is confidently too small.
A finite variance is the second, and its failure is more dramatic. The Cauchy distribution, which describes the horizontal position where a randomly angled beam from a point source meets a line, has no finite mean or variance. The average of independent Cauchy variables has exactly the same distribution as a single one, no matter how large is. Averaging accomplishes literally nothing, and no amount of data helps. Real heavy-tailed data, in insurance and finance, sits between this extreme and the well-behaved case, and it converges slowly enough that the normal approximation misleads at the sample sizes people actually have.
The third is a matter of rate rather than of validity. The theorem says the limit is normal, and says nothing about how large must be for the approximation to be usable. For a symmetric ingredient like a die, ten is plenty. For a heavily skewed one, hundreds or thousands may not be, and the tails converge far more slowly than the centre, which is precisely where the answer usually matters.
The machinery of the subject is now complete: sample spaces, conditioning, inversion, quantities, spread, the standard laws, and the limit theorems that connect them back to observation. What remains is the reliable business of applying it wrongly, which is the last lesson.