Predicting what a player will do usually requires knowing what they expect everyone else to do, and there is one case where it does not: an option that is worse than another whatever happens.
An option no belief can justify
The previous lesson left two airlines with a grid and no prediction. Each sets a fare of £120 or £90, and in thousands of pounds the payoffs are 9.0, 9.0 when both hold high, 6.0, 6.0 when both cut, and 10.2, 2.7 in favour of whichever one undercuts alone.
Look at the grid from A's side, one column at a time. If B holds at £120, A earns 9.0 by holding and 10.2 by cutting, so cutting is better. If B cuts to £90, A earns 2.7 by holding and 6.0 by cutting, so cutting is better again. A does not need to know, guess or estimate what B will do. Whatever B does, cutting pays more.
That is strict dominance. Strategy strictly dominates for player when
meaning for every combination of choices the other players might make. A rational player never plays a strictly dominated strategy, and the argument is airtight in a way that almost nothing else in this course is: whatever beliefs the player holds, however wrong those beliefs are, the dominating strategy yields a higher expected payoff, because it yields more against every single thing the others could do. The converse holds too, in a form worth remembering: in a finite two-player game, a strategy is strictly dominated exactly when there is no belief at all about the opponent to which it is a best reply.
The grid is symmetric, so the same argument runs for B. Cutting is dominant for both, both cut, and the prediction is 6.0, 6.0. Both firms end up £3,000 worse off than in the cell they were both trying to reach, and no one made a mistake.
The prisoner's dilemma
That structure has a name, and it is the most reproduced object in the social sciences. Merrill Flood and Melvin Dresher constructed the game at the RAND Corporation in January 1950 and ran it as a hundred-round experiment between two colleagues, who cooperated in about sixty of the hundred rounds and left irritated notes about each other's play. Albert Tucker, needing to explain the payoff structure to an audience of psychologists later that year, dressed it in the story that stuck.
Two suspects are held separately. The evidence supports a minor charge with a one-year sentence. Each is offered the same deal: testify against the other and walk free, while the other serves ten years. If both testify, the deal is withdrawn from both and each serves six years. Payoffs are negative years served, so the grid is:
| Prisoner 1 | 2 stays silent | 2 testifies |
|---|---|---|
| Stays silent | -1, -1 | -10, 0 |
| Testifies | 0, -10 | -6, -6 |
Testifying strictly dominates silence: 0 beats -1 in the left column, -6 beats -10 in the right. Both testify, both serve six years, and both would have served one had they held. The outcome is Pareto dominated, meaning there is another outcome both players strictly prefer, and rational individual choice walks straight past it.
The general form uses four letters. Write for the temptation payoff to the sole defector, for the reward when both cooperate, for the punishment when both defect, and for the sucker's payoff when you cooperate alone. The game is a prisoner's dilemma exactly when
and repeated-play analysis normally adds a second condition, , which says that mutual cooperation beats taking turns at exploiting each other. Both conditions are checkable arithmetic, and a great many situations described as prisoner's dilemmas in print fail one of them.
Real instances are everywhere, and they are not moral parables. Two OPEC members each gain by producing above quota while the other holds back, and the cartel price collapses when both do. Two nations each gain by arming while the other disarms. A cyclist gains by doping while the peloton stays clean, and the sport ends up with a doped peloton riding at the same relative speeds. A farmer gains by drawing from the shared aquifer, and the aquifer runs dry. In each the dominant strategy is individually correct and collectively ruinous, and the standard repairs, which are contracts, regulators, monitoring, and above all repetition, all work by changing the payoffs rather than by improving anyone's reasoning.
Example. Two firms choose to cooperate or defect. Mutual cooperation pays 4 each, mutual defection pays 2 each, and a sole defector gets 6 while the cooperator gets 1. Is this a prisoner's dilemma?
Read off , , , . The ordering condition holds: . The second condition also holds, since exceeds . It is a prisoner's dilemma in the full sense, and defection strictly dominates for both, since 6 beats 4 against a cooperator and 2 beats 1 against a defector.
Now you. Same structure with mutual cooperation paying 3 each, mutual defection paying 1 each, a sole defector getting 7 and the cooperator getting 0. Does this game satisfy both conditions?
Answer
, so the ordering condition holds and defection is still strictly dominant. But is less than , so the second condition fails. Two players who could coordinate would do better taking turns at exploitation, averaging 3.5 each, than by cooperating every round. That matters once repetition is on the table, because a cooperative arrangement is not the best thing the pair can arrange.
Iterated deletion
Dominance is worth much more than the games in which someone has a dominant strategy, because deleting a dominated strategy leaves a smaller game in which something else may now be dominated.
Take a three-by-three game. Player 1 chooses the row from U, M and D, player 2 chooses the column from L, C and R, and the cells hold the pair of payoffs:
| Player 1 | L | C | R |
|---|---|---|---|
| U | 4, 3 | 5, 1 | 6, 2 |
| M | 2, 1 | 8, 4 | 3, 6 |
| D | 3, 0 | 9, 6 | 2, 8 |
No player has a dominant strategy here. Start with player 2, whose payoffs are the second number in each cell. Compare column C with column R: C pays 1 against U where R pays 2, C pays 4 against M where R pays 6, C pays 6 against D where R pays 8. R is better in every row, so C is strictly dominated and player 2 will not choose it. Delete the C column.
Now look again at player 1, who is left choosing among rows against L and R only. Against L, U pays 4, M pays 2 and D pays 3; against R, U pays 6, M pays 3 and D pays 2. U beats both of the others in both columns, so M and D are now strictly dominated. Note that they were not dominated in the original game: D paid 9 against C, better than U's 5. It is the deletion of C, itself justified only by player 2's rationality, that makes them dominated. Delete M and D.
Player 2 is now choosing against U alone, where L pays 3 and R pays 2. Delete R. One cell survives: (U, L), paying 4 to player 1 and 3 to player 2.
The procedure is called iterated deletion of strictly dominated strategies, and its assumptions escalate at every round. Deleting C requires only that player 2 is rational. Deleting M and D requires that player 1 is rational and knows that player 2 is. Deleting R requires that player 2 knows that player 1 knows that player 2 is rational. That is the common knowledge assumption from the previous lesson being spent, one level per round, which is why the deeper rounds of these arguments should be trusted less than the first.
One reassurance about the procedure: for strict dominance, the order of deletion does not matter. Whatever sequence of legal deletions you make, the same set of strategies survives, because a strategy strictly dominated in the full game is still strictly dominated in any subgame you reach by deleting other strategies. That is not a small property, and the next section shows how badly it fails when "strictly" is weakened.
Example. Apply iterated deletion to this game.
| Player 1 | L | C | R |
|---|---|---|---|
| T | 4, 2 | 2, 3 | 3, 1 |
| M | 5, 1 | 3, 4 | 2, 2 |
| B | 2, 0 | 1, 2 | 1, 3 |
Player 1's row B pays 2, 1, 1 against the three columns, while T pays 4, 2, 3. T beats B everywhere, so B goes. With rows T and M left, player 2's column C pays 3 and 4, column L pays 2 and 1, column R pays 1 and 2, so C strictly dominates both and L and R go. Player 1 then compares T's 2 with M's 3 in column C and takes M. The prediction is (M, C), paying 3 to player 1 and 4 to player 2.
Now you. In the three-by-three game solved in the section above, which single deletion has to come first, and what would go wrong if player 1 tried to delete M at the very start?
Answer
Only column C is dominated in the original game, so it must go first. Deleting M at the start is illegitimate because M is not dominated there: against C it pays 8 while U pays 5, so M is better than U in that column. Player 1 can only rule M out after ruling out C, which is a claim about player 2's rationality rather than about the payoff table alone.
Weak dominance, and why it is treated with suspicion
Strategy weakly dominates when it is at least as good against everything and strictly better against something. That sounds like a harmless relaxation. It is not.
| Player 1 | L | R |
|---|---|---|
| T | 1, 1 | 0, 0 |
| M | 1, 1 | 2, 1 |
| B | 0, 0 | 2, 1 |
For player 1, M weakly dominates T, since it ties at L and pays 2 against 0 at R. M also weakly dominates B, since it pays 1 against 0 at L and ties at R. Both deletions are legal, and they lead to different places.
Delete T first. Player 2, choosing between L and R against rows M and B, finds that R ties with L against M and beats it against B, so L is weakly dominated and goes. What survives pays 2 to player 1 and 1 to player 2. Now start again and delete B first instead. Player 2, choosing against rows T and M, finds that L ties with R against M and beats it against T, so R goes. What survives pays 1 to player 1 and 1 to player 2. Same game, two legal orders, and player 1's payoff is 2 or 1 depending on which was chosen.
So iterated weak dominance is not a well-defined solution procedure, and results derived with it should always say which order was used. Weak dominance is still worth having as a one-shot argument about a single strategy, and one of the most useful results in the whole subject is of that kind: in a sealed-bid second-price auction, bidding your true value weakly dominates every other bid. That argument, and the auction it lives in, arrive near the end of this course.
Dominance by a mixed strategy
A strategy can fail to be dominated by any single alternative and still be dominated by a randomisation over alternatives. Since the payoffs are utilities, the expected payoff of a randomisation is a legitimate comparison, which is exactly what the utility axioms from the previous lesson were for.
| Player 1 | L | R |
|---|---|---|
| U | 4, 0 | 0, 0 |
| M | 1, 0 | 1, 0 |
| D | 0, 0 | 4, 0 |
M is not dominated by U, which collapses to 0 against R, nor by D, which collapses to 0 against L. But consider the strategy "play U with probability 0.5 and D with probability 0.5". Against L it yields ; against R it yields . Two beats one in both columns, so M is strictly dominated by the mixture and can be deleted, even though a player using it would never write down a plan that includes M's guaranteed 1.
Example. Replace M's payoffs above with 3 against L and 3 against R, leaving U and D as they were. Is M dominated by any mixture of U and D?
A mixture playing U with probability pays against L and against R. To dominate M it must beat 3 in both columns, so and , meaning and at once. No such exists, so M survives. Its guaranteed 3 is too good to be beaten by a gamble on the extremes.
Now you. Now let U pay 5 against L and 0 against R, D pay 0 against L and 5 against R, and M pay 2 against both. Is M strictly dominated by a mixture?
Answer
The mixture pays against L and against R, and beating 2 in both columns needs and . Any strictly between 0.4 and 0.6 works, and gives 2.5 against either column, so M is strictly dominated and can be deleted.
What dominance cannot do
For all its rigour, dominance answers almost nothing. Most games have no dominated strategies at all, and iterated deletion in those games removes nothing and predicts nothing. Even when it bites, it often leaves a large rectangle of survivors rather than a single cell. It is a filter, not a solution concept, and a subject that stopped here would be able to analyse the prisoner's dilemma and very little else.
The empirical record is mixed even where the theory is sharpest. First plays of a one-shot prisoner's dilemma in the laboratory produce cooperation in something like a third to a half of cases, well above the zero the dominance argument predicts, and the rate declines but does not vanish with experience. People also play strictly dominated strategies in games where the dominance is hidden behind a step of arithmetic, which is evidence about attention rather than about preferences. Both facts return in the final lesson.
What is needed is a weaker requirement that still cuts: not "better whatever they do", which is rare, but "best given what they are actually doing". Making that idea non-circular, when what they are doing depends in turn on what you are doing, is the single most important step in the subject, and it is next.