Everything from here on is a contest between hypotheses that nobody can settle by looking. Did the universe have a cause? Is the fine structure of physics surprising? Is the silence of the sky evidence that we are alone? These cannot be argued with the tools of lesson one alone, because the answer is never that one side knows and the other does not. What is needed is a way of saying how much a piece of evidence should move you, and that is this lesson.
The probability course develops this properly. Here we need three ideas and one formula.
Belief comes in degrees
The first move is to stop treating belief as a switch. You do not merely believe or disbelieve that it will rain tomorrow, you have some level of confidence, and you act on it by carrying an umbrella or not. Philosophers call that level a credence, a number between 0 and 1.
This matters because most bad arguments in the rest of the course exploit the binary picture. "You cannot prove there is no God, so atheism is faith too" works only if belief has two settings. Once you allow degrees, the reply is easy: the claim is not that the probability is zero, it is that it is low, and a low credence is a different thing from a high one even though neither is certainty. Lesson one killed the certainty confusion; this kills its more resilient cousin.
The second move is to notice that a credence has to start somewhere before the evidence arrives. That starting point is the prior, and it is where most real disagreement lives. Two people who agree completely about the evidence can end up far apart because they began far apart, and pretending otherwise produces arguments that go nowhere for hours.
Example. A test for a disease is 99 percent accurate in both directions. The disease affects 1 person in 10,000. You test positive. How worried should you be?
Much less than the number suggests. Imagine a million people. About 100 have the disease and 99 of them test positive. The other 999,900 are healthy, and 1 percent of them, about 9,999, test positive anyway. So roughly 10,098 people test positive and about 99 of them are ill: just under 1 percent. The test is excellent and you are still very probably fine, because the prior was so low that even strong evidence cannot drag it far. This is the single most useful pattern in the lesson. When someone presents evidence for a claim that starts out very improbable, the evidence has to be extraordinarily discriminating to make the claim likely, and "the test is 99 percent accurate" is nowhere near enough.
Now you. Someone replies that the calculation is a trick, because your particular test came back positive and the other 999,999 people are irrelevant to your case. Where does this go wrong?
Answer
It mistakes what the base rate is doing. Nobody is claiming the other people affect your biology, only that they tell you what kind of positive result yours is likely to be. A positive result is a category with two sorts of member, true and false positives, and the base rate determines how that category is composed. Asking how worried to be just is asking which sort yours probably is, and that question cannot be answered from the accuracy of the test alone. The mistake has a name, base rate neglect, it is extremely common among trained professionals as well as everyone else, and it recurs in lesson thirteen when people reason about how surprising it is that the universe permits life.
Bayes in one example
The formula that makes this exact is Bayes' theorem. Write for your prior in a hypothesis, for how likely the evidence is if the hypothesis is true, and for how likely the evidence is overall. Then
Read it as an instruction rather than an equation. Your new confidence is your old confidence, multiplied by how well the hypothesis predicted what you saw, divided by how likely that observation was anyway.
The last part is what people forget, and it is the whole game. Evidence that any hypothesis would have predicted is worthless, because the denominator is as large as the numerator and nothing moves. Evidence is only evidence to the extent that it was more expected on one hypothesis than on its rivals.
That gives the practical test, and it is worth memorising because the rest of the course uses it constantly. Do not ask whether the evidence fits the hypothesis. Ask how much more likely the evidence is if the hypothesis is true than if it is false. That ratio, the likelihood ratio, is the strength of the evidence, and everything else is presentation.
Example. An advocate says the fine structure of physics is evidence for design, because a designed universe would certainly permit life. Apply the test.
The claim gives you the numerator, that is high: if there is a designer who wants life, a life-permitting universe is very likely. It says nothing about the denominator, and the denominator is where the argument lives. To make this evidence you must also say that a life-permitting universe is unlikely otherwise, and defending that requires a view about how physical constants could have been, which is a hard and contested question rather than an obvious one. The argument as stated is incomplete rather than wrong, and lesson ten takes up the missing half. What the test does immediately is show you exactly which half is missing, which is why it is worth having before meeting the argument in its full dress.
Now you. A different advocate argues that the universe is not evidence for design, because whatever the constants had been, someone would have found them remarkable. Is that a good reply?
Answer
It is a real point badly aimed. The observation that any specific outcome is improbable is true and does not by itself defeat the design argument, because the design argument does not rest on the constants being specific, it rests on their falling in a narrow range that permits observers at all. Being dealt a particular bridge hand is improbable and unremarkable; being dealt thirteen spades is equally improbable and calls for explanation, because it belongs to a tiny and independently interesting class. The reply is therefore too quick as stated. Its real force appears in lesson thirteen, where the relevant class turns out to be selected by the fact that we are here to notice, and that is an argument about the denominator rather than about improbability in general.
Why unfalsifiable claims get nothing
Now a consequence that will do a lot of work later.
Suppose someone offers a hypothesis so flexible that it accommodates any observation whatever. Whatever happens, the theory can explain it afterwards. People usually treat this as a strength, and say the theory has never been refuted.
Apply the test. If the hypothesis predicts every possible observation equally, then is the same for all E, and the likelihood ratio against any rival is 1. No observation discriminates. The theory is never refuted and never confirmed either, and these are the same fact rather than two facts. Its survival is not a track record, it is an artefact of its shape.
This is the precise version of the point Karl Popper made with falsifiability, and the Bayesian form is more useful, since it comes in degrees rather than being a pass or fail. A theory can be a little bit accommodating and lose a little bit of confirming power. Lesson nine takes this further and applies it to whole pictures of reality.
Example. Two people argue about a claimed psychic. The psychic performs well under casual conditions and poorly under controlled ones. The believer says the laboratory atmosphere of scepticism disrupts the ability. What is wrong with that move, in the language of this lesson?
It converts a discriminating experiment into a non-discriminating one after the result is known. Before the test, the psychic hypothesis predicted success and the fraud hypothesis predicted failure, so the likelihood ratio was large and the experiment was worth running. The rescue adds a clause saying failure is also expected, which makes both outcomes compatible with the hypothesis and drives the ratio to 1. Notice the timing is the offence rather than the content: if the disruption claim had been stated in advance, with a specification of what counts as a disruptive atmosphere, it would have been a testable refinement. Added afterwards, it converts a theory that could have earned support into one that cannot. Lesson nine calls this ad hoc rescue.
Now you. A friend says: "The simulation hypothesis is unfalsifiable, so it is unscientific, so it is false." Which step is wrong?
Answer
The last one. Unfalsifiability is a defect in the relation between a hypothesis and evidence, not a property of the world, and it gives you no licence to conclude the hypothesis is false. What follows is that the evidence cannot support it, so your credence should stay wherever your prior put it. That may be low, for reasons of parsimony that lesson nine sets out, but the lowness comes from the prior rather than from any observation. The distinction matters because the same bad step is made in both directions in the God lessons: "you cannot test it, therefore it is false" is no better than "you cannot refute it, therefore it is true", and both are failures to see that unfalsifiability leaves you exactly where you started.
What this does not give you
Two honest limitations, since the machinery is often oversold.
It does not supply priors. Bayes tells you how to update, not where to begin, and in the arguments ahead the priors are frequently the whole dispute. Someone who assigns a very low prior to any supernatural claim and someone who assigns a moderate one will read the same evidence and reach different places, both correctly by their own lights. What the framework does is make that disagreement visible and locate it, which is more than most arguments about God ever manage.
It also does not make the numbers real. Nobody has a well-defined credence in the existence of a deity to three decimal places, and writing one down does not create one. The value of the apparatus in a subject like this is qualitative: it tells you which direction the evidence pushes, roughly how hard, and when a piece of evidence is doing nothing at all. That is enough to settle a surprising number of arguments.
The tools are now in place. Before using them on the hardest questions, one more preparation is needed, because the errors in lessons ten to fourteen are not mainly errors of logic. They are errors of perspective, made by a particular kind of animal reasoning about its own importance.