Uncertainty is not the same thing as ignorance, because some uncertain things are more uncertain than others, and that comparison is what a number is for. This lesson asks what such a number could possibly mean, gets two incompatible answers, and finds that both answers obey the same rules. Nothing here needs more than arithmetic and the algebra of the foundations course.
The rule that started it
The oldest working definition of probability is a ratio of counts. List every case the situation can produce, decide which of them count as the event you care about, and divide:
Gerolamo Cardano wrote this down in his Liber de Ludo Aleae around 1564, sitting on the manuscript until it was printed in 1663, and Pierre-Simon Laplace made it the formal definition in 1812 with one crucial rider: the cases must be equally possible. With that rider the rule works. A fair die has six cases, one of which is a six, so . Two dice have cases, of which six sum to seven, so as well, while only five sum to eight, giving .
That last pair of numbers is already worth something, because it explains an observation gamblers had made for centuries without being able to justify: seven comes up more often than eight, even though both can be made in several ways. The counting has to be done over the ordered pairs, not over the unordered totals, and the whole history of early probability is people getting that distinction wrong.
Two games that differ by a hair
In 1654 Antoine Gombaud, who styled himself the Chevalier de Méré, put a complaint to Blaise Pascal. He made money betting that at least one six would appear in four rolls of a single die, and lost money betting that at least one double six would appear in twenty-four rolls of two dice. His reasoning said the two should be identical: one chance in six, four rolls, and one chance in thirty-six, twenty-four rolls, both giving four sixths of a certainty. His purse said otherwise.
The counting rule settles it, provided you count the failures rather than the successes. Four rolls of one die produce equally likely sequences, and of them contain no six, so
Twenty-four rolls of a pair produce sequences, of which avoid the double six, so
One game is a shade better than even and the other a shade worse. The gap is , about one bet in thirty-eight, which is invisible over an evening and ruinous over a career. De Méré had played enough to feel a difference of two and a half percentage points. The correspondence between Pascal and Fermat that followed this question is where the subject begins.
Example. What is the probability of getting at least one six in five rolls of a fair die?
Count the failures. Each roll avoids a six in five of its six cases, and the rolls are laid out as a sequence, so of the sequences contain no six at all. The complement gives .
Now you. What is the probability of getting at least one six in three rolls of a fair die?
Answer
, so three rolls is a losing bet and four rolls a winning one.
Where counting cases runs out
The classical rule has two faults, and neither is repairable from inside it.
The first is circularity. "Equally possible" means equally probable, so the definition of probability contains the word it is defining. For a die this is harmless, because the symmetry of a cube gives an independent reason to believe the six cases are interchangeable: relabel the faces and nothing physical changes. Away from manufactured symmetric objects the escape closes. A drawing pin tossed on a table lands point up or point down, which is two cases, and nobody believes the answer is . Nothing about the pin says the two cases are interchangeable, and the only way to find out is to throw it.
The second fault is that the rule is silent whenever the cases cannot be made symmetric at all. What is the probability that a particular patient survives five years, that a bridge design fails, that it rains tomorrow? There is no bag of equally possible cases to count. These are the questions people actually have, and the classical definition does not so much answer them badly as refuse to accept them.
So the counting rule is not a definition. It is a correct calculation for one special structure, a finite space of interchangeable outcomes, and the general notion has to come from somewhere else.
What the record actually shows
The obvious somewhere else is measurement. If the probability of heads is the fraction of heads in the long run, then the way to find it is to toss the coin many times. People have done exactly this. Georges-Louis Leclerc, Comte de Buffon, recorded heads in tosses. Karl Pearson recorded in . The most careful record is John Kerrich's, made while interned in Denmark during the Second World War and published in 1946, and it is worth seeing in full because it shows two things at once.
| Tosses | Heads | Heads minus half the tosses | Fraction heads |
|---|---|---|---|
| 10 | 4 | ||
| 100 | 44 | ||
| 1000 | 502 | ||
| 5000 | 2533 | ||
| 10000 | 5067 |
Read the third column and the coin looks worse and worse: Kerrich ends heads clear of half, having been only clear at a hundred tosses. Read the fourth and it looks better and better, closing on from . Both columns are correct, and they are not in conflict. The absolute surplus of heads grows without limit while the surplus divided by the number of tosses shrinks towards nothing. Whichever of these you call "the coin settling down" determines what you think probability is, and the later lesson on the laws of large numbers gives the exact rate at which each happens.
The fatal difficulty with defining probability as the long-run fraction is in the phrase "long run". No finite record ever hands you a number: Kerrich's fraction was , not , and another ten thousand tosses would have given a third value. To say the fraction tends to a limit is to make a claim no experiment can confirm, and worse, it is a claim that is not even guaranteed by the theory, since a fair coin can in principle give heads forever. Frequency tells you how to estimate a probability. It does not tell you what one is.
Example. Kerrich was heads below half at a hundred tosses and above at ten thousand. In which record was the coin behaving more like a fair one?
Compare fractions, not counts. At a hundred tosses the fraction was , which misses by . At ten thousand it was , missing by , roughly nine times closer. The larger discrepancy in raw heads belongs to the far larger experiment, and it is the fraction that carries the information about the coin.
Now you. At one thousand tosses Kerrich had heads. Give the fraction and its distance from , and say how it compares with the ten thousand toss figure.
Answer
, which misses by , closer than the at ten thousand tosses. Convergence is not tidy: a fraction can drift back out again, and does.
Probability as a price
The third reading abandons repetition altogether. On this view a probability is a degree of belief, and it is measured the way beliefs have always been measured in practice, by what someone will bet. Frank Ramsey in 1926 and Bruno de Finetti in 1931 made this precise. If you regard a payout of one unit on event as worth exactly units to buy or to sell, then is your probability for .
This sounds like an invitation to say anything, and it is not, because inconsistent prices can be robbed. Bookmakers quote odds "against": at to against, a winning stake of returns , which corresponds to a probability of . Suppose a race with three horses is priced at to , to and to against. The implied probabilities are , and , and they sum to . Back all three, staking exactly those fractions of a unit. Whichever horse wins, the return is exactly , because . Total outlay , guaranteed return , guaranteed profit per unit staked out, a return of percent on money with no risk whatsoever.
A set of prices that permits this is called a Dutch book, and the theorem that Ramsey and de Finetti proved is that you are safe from one if and only if your prices lie between and and add to exactly across a set of alternatives exactly one of which must happen. Coherence, not frequency, forces the arithmetic. Real bookmakers make the sum exceed on purpose, and the excess is their margin.
Example. A market prices three outcomes at evens ( to ), to and to against. Is it coherent, and if not, what is the risk-free return?
The implied probabilities are , and , summing to , which is less than , so the market is incoherent in the punter's favour. Staking , and returns exactly whichever outcome occurs, for an outlay of : a profit of per unit returned, which is percent on the money staked.
Now you. A two-outcome market is priced at to against on each side. What do the implied probabilities sum to, and who can be robbed?
Answer
Each price implies , so the sum is . Staking on each costs and returns for certain, a guaranteed profit of , so the bookmaker is the one being robbed.
What both readings share
The frequentist reads a probability as a fact about a repeatable setup, and refuses to assign one to a unique event. The Bayesian reads it as a coherent price, and assigns one to anything. They disagree about what the number means, and about which questions are legitimate. They do not disagree about the arithmetic, and this is the fact the rest of the subject is built on.
Both give every probability a value in , because a fraction of a total cannot be negative or exceed the whole, and a price outside that range is a Dutch book in one line. Both make the probabilities of a set of mutually exclusive alternatives, one of which must occur, sum to exactly : the counting rule because the favourable cases partition the total, the betting rule because coherence demands it. Both make the probability of " or " equal to the sum of the separate probabilities when and cannot both happen, since disjoint sets of cases add and so do the stakes that cover them.
That is three rules, and they were reached three times over from unrelated starting points. Andrey Kolmogorov's move in 1933 was to stop asking what probability is and take those three statements as axioms, with everything else to be proved from them. It is the same move that made group theory out of symmetry and metric spaces out of distance, and it is why probability became a branch of mathematics rather than a collection of gambling results.
What is still missing
Taking the rules as axioms leaves an obvious gap. Rules have to be rules about something, and every statement above quietly assumed a background list: the six faces, the sequences, the three horses. "The probability of or " only means anything once and are the kind of object that can be combined with "or" at all.
So the first construction the subject needs is not a formula but a set: the collection of everything that could happen, with events living inside it as subsets, and "or", "and" and "not" becoming union, intersection and complement. That is the next lesson, and it turns the three rules above into three axioms with enough structure that real theorems follow from them.