Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Measuring uncertainty

Uncertainty is not the same thing as ignorance, because some uncertain things are more uncertain than others, and that comparison is what a number is for. This lesson asks what such a number could possibly mean, gets two incompatible answers, and finds that both answers obey the same rules. Nothing here needs more than arithmetic and the algebra of the foundations course.

The rule that started it

The oldest working definition of probability is a ratio of counts. List every case the situation can produce, decide which of them count as the event you care about, and divide:

P(A)=number of cases favourable to Anumber of cases in total

Gerolamo Cardano wrote this down in his Liber de Ludo Aleae around 1564, sitting on the manuscript until it was printed in 1663, and Pierre-Simon Laplace made it the formal definition in 1812 with one crucial rider: the cases must be equally possible. With that rider the rule works. A fair die has six cases, one of which is a six, so P(six)=1/60.1667. Two dice have 6×6=36 cases, of which six sum to seven, so P(sum=7)=6/360.1667 as well, while only five sum to eight, giving 5/360.1389.

That last pair of numbers is already worth something, because it explains an observation gamblers had made for centuries without being able to justify: seven comes up more often than eight, even though both can be made in several ways. The counting has to be done over the 36 ordered pairs, not over the 21 unordered totals, and the whole history of early probability is people getting that distinction wrong.

Two games that differ by a hair

In 1654 Antoine Gombaud, who styled himself the Chevalier de Méré, put a complaint to Blaise Pascal. He made money betting that at least one six would appear in four rolls of a single die, and lost money betting that at least one double six would appear in twenty-four rolls of two dice. His reasoning said the two should be identical: one chance in six, four rolls, and one chance in thirty-six, twenty-four rolls, both giving four sixths of a certainty. His purse said otherwise.

The counting rule settles it, provided you count the failures rather than the successes. Four rolls of one die produce 64=1296 equally likely sequences, and 54=625 of them contain no six, so

P(at least one six)=1-6251296=67112960.5177

Twenty-four rolls of a pair produce 3624 sequences, of which 3524 avoid the double six, so

P(at least one double six)=1-(3536)240.4914

One game is a shade better than even and the other a shade worse. The gap is 0.0263, about one bet in thirty-eight, which is invisible over an evening and ruinous over a career. De Méré had played enough to feel a difference of two and a half percentage points. The correspondence between Pascal and Fermat that followed this question is where the subject begins.

Example. What is the probability of getting at least one six in five rolls of a fair die?

Count the failures. Each roll avoids a six in five of its six cases, and the rolls are laid out as a sequence, so 55=3125 of the 65=7776 sequences contain no six at all. The complement gives 1-3125/7776=4651/77760.5981.

Now you. What is the probability of getting at least one six in three rolls of a fair die?

Answer

1-(5/6)3=1-125/216=91/2160.4213, so three rolls is a losing bet and four rolls a winning one.

Where counting cases runs out

The classical rule has two faults, and neither is repairable from inside it.

The first is circularity. "Equally possible" means equally probable, so the definition of probability contains the word it is defining. For a die this is harmless, because the symmetry of a cube gives an independent reason to believe the six cases are interchangeable: relabel the faces and nothing physical changes. Away from manufactured symmetric objects the escape closes. A drawing pin tossed on a table lands point up or point down, which is two cases, and nobody believes the answer is 1/2. Nothing about the pin says the two cases are interchangeable, and the only way to find out is to throw it.

The second fault is that the rule is silent whenever the cases cannot be made symmetric at all. What is the probability that a particular patient survives five years, that a bridge design fails, that it rains tomorrow? There is no bag of equally possible cases to count. These are the questions people actually have, and the classical definition does not so much answer them badly as refuse to accept them.

So the counting rule is not a definition. It is a correct calculation for one special structure, a finite space of interchangeable outcomes, and the general notion has to come from somewhere else.

What the record actually shows

The obvious somewhere else is measurement. If the probability of heads is the fraction of heads in the long run, then the way to find it is to toss the coin many times. People have done exactly this. Georges-Louis Leclerc, Comte de Buffon, recorded 2048 heads in 4040 tosses. Karl Pearson recorded 12012 in 24000. The most careful record is John Kerrich's, made while interned in Denmark during the Second World War and published in 1946, and it is worth seeing in full because it shows two things at once.

TossesHeadsHeads minus half the tossesFraction heads
104-10.4000
10044-60.4400
1000502+20.5020
50002533+330.5066
100005067+670.5067

Read the third column and the coin looks worse and worse: Kerrich ends 67 heads clear of half, having been only 6 clear at a hundred tosses. Read the fourth and it looks better and better, closing on 0.5 from 0.44. Both columns are correct, and they are not in conflict. The absolute surplus of heads grows without limit while the surplus divided by the number of tosses shrinks towards nothing. Whichever of these you call "the coin settling down" determines what you think probability is, and the later lesson on the laws of large numbers gives the exact rate at which each happens.

The fatal difficulty with defining probability as the long-run fraction is in the phrase "long run". No finite record ever hands you a number: Kerrich's fraction was 0.5067, not 0.5, and another ten thousand tosses would have given a third value. To say the fraction tends to a limit is to make a claim no experiment can confirm, and worse, it is a claim that is not even guaranteed by the theory, since a fair coin can in principle give heads forever. Frequency tells you how to estimate a probability. It does not tell you what one is.

Example. Kerrich was 6 heads below half at a hundred tosses and 67 above at ten thousand. In which record was the coin behaving more like a fair one?

Compare fractions, not counts. At a hundred tosses the fraction was 44/100=0.44, which misses 0.5 by 0.06. At ten thousand it was 5067/10000=0.5067, missing by 0.0067, roughly nine times closer. The larger discrepancy in raw heads belongs to the far larger experiment, and it is the fraction that carries the information about the coin.

Now you. At one thousand tosses Kerrich had 502 heads. Give the fraction and its distance from 0.5, and say how it compares with the ten thousand toss figure.

Answer

502/1000=0.502, which misses 0.5 by 0.002, closer than the 0.0067 at ten thousand tosses. Convergence is not tidy: a fraction can drift back out again, and does.

Probability as a price

The third reading abandons repetition altogether. On this view a probability is a degree of belief, and it is measured the way beliefs have always been measured in practice, by what someone will bet. Frank Ramsey in 1926 and Bruno de Finetti in 1931 made this precise. If you regard a payout of one unit on event A as worth exactly p units to buy or to sell, then p is your probability for A.

This sounds like an invitation to say anything, and it is not, because inconsistent prices can be robbed. Bookmakers quote odds "against": at 3 to 1 against, a winning stake of 1 returns 4, which corresponds to a probability of 1/(3+1)=0.25. Suppose a race with three horses is priced at 2 to 1, 3 to 1 and 4 to 1 against. The implied probabilities are 1/3, 1/4 and 1/5, and they sum to 0.7833. Back all three, staking exactly those fractions of a unit. Whichever horse wins, the return is exactly 1, because 3×13=4×14=5×15=1. Total outlay 0.7833, guaranteed return 1, guaranteed profit 0.2167 per unit staked out, a return of 27.7 percent on money with no risk whatsoever.

A set of prices that permits this is called a Dutch book, and the theorem that Ramsey and de Finetti proved is that you are safe from one if and only if your prices lie between 0 and 1 and add to exactly 1 across a set of alternatives exactly one of which must happen. Coherence, not frequency, forces the arithmetic. Real bookmakers make the sum exceed 1 on purpose, and the excess is their margin.

Example. A market prices three outcomes at evens (1 to 1), 5 to 1 and 9 to 1 against. Is it coherent, and if not, what is the risk-free return?

The implied probabilities are 1/2, 1/6 and 1/10, summing to 0.7667, which is less than 1, so the market is incoherent in the punter's favour. Staking 0.5, 0.1667 and 0.1 returns exactly 1 whichever outcome occurs, for an outlay of 0.7667: a profit of 0.2333 per unit returned, which is 30.4 percent on the money staked.

Now you. A two-outcome market is priced at 2 to 1 against on each side. What do the implied probabilities sum to, and who can be robbed?

Answer

Each price implies 1/3, so the sum is 2/3. Staking 1/3 on each costs 0.6667 and returns 1 for certain, a guaranteed profit of 0.3333, so the bookmaker is the one being robbed.

What both readings share

The frequentist reads a probability as a fact about a repeatable setup, and refuses to assign one to a unique event. The Bayesian reads it as a coherent price, and assigns one to anything. They disagree about what the number means, and about which questions are legitimate. They do not disagree about the arithmetic, and this is the fact the rest of the subject is built on.

Both give every probability a value in [0,1], because a fraction of a total cannot be negative or exceed the whole, and a price outside that range is a Dutch book in one line. Both make the probabilities of a set of mutually exclusive alternatives, one of which must occur, sum to exactly 1: the counting rule because the favourable cases partition the total, the betting rule because coherence demands it. Both make the probability of "A or B" equal to the sum of the separate probabilities when A and B cannot both happen, since disjoint sets of cases add and so do the stakes that cover them.

That is three rules, and they were reached three times over from unrelated starting points. Andrey Kolmogorov's move in 1933 was to stop asking what probability is and take those three statements as axioms, with everything else to be proved from them. It is the same move that made group theory out of symmetry and metric spaces out of distance, and it is why probability became a branch of mathematics rather than a collection of gambling results.

What is still missing

Taking the rules as axioms leaves an obvious gap. Rules have to be rules about something, and every statement above quietly assumed a background list: the six faces, the 1296 sequences, the three horses. "The probability of A or B" only means anything once A and B are the kind of object that can be combined with "or" at all.

So the first construction the subject needs is not a formula but a set: the collection of everything that could happen, with events living inside it as subsets, and "or", "and" and "not" becoming union, intersection and complement. That is the next lesson, and it turns the three rules above into three axioms with enough structure that real theorems follow from them.