Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Game Theory

Strategy when the best move depends on someone else's: dominance and equilibrium, repetition and reputation, bargaining, and where the predictions fail.

What a game is

A decision problem becomes a game the moment the best thing for you to do depends on what somebody else decides at the same time, and the whole subject exists because that dependence breaks ordinary decision-making.

Against nature, against a person

Deciding whether to carry an umbrella is a problem with one decision-maker. The weather has no interest in whether you get wet, so you can attach a probability to rain, weigh the cost of carrying the umbrella against the cost of a soaking, and take the option with the better expected payoff. That is decision theory, and the machinery for it is the probability course: a distribution over states of the world, a payoff for each combination of state and action, and an expectation to maximise.

Now change one thing. Suppose the rain is produced by somebody who is paid whenever you get wet, and who knows you are reasoning about umbrellas. There is no fixed probability of rain to look up, because the chance of rain depends on the chance you carry an umbrella, which depends on what you think the chance of rain is. Each side's calculation is an input to the other's. The regress is real, not a trick of presentation, and no amount of care about probabilities dissolves it.

Game theory is the machinery for that second case. It was assembled for the first time in John von Neumann and Oskar Morgenstern's Theory of Games and Economic Behavior in 1944, on foundations von Neumann had laid in a 1928 paper on parlour games, and the word "game" has been misleading people ever since. The subject is applied to price wars, arms control, evolution, auctions, plea bargains, divorce settlements and the timing of an election. What makes something a game is not that it is played for fun but that the payoff to your choice is not yours alone to determine.

The three ingredients

Strip the story away and every strategic situation has the same skeleton, and it has exactly three parts.

There is a set of players, indexed i=1,2,,n, meaning the decision-makers whose choices matter. There is, for each player, a set of strategies Si, meaning the options that player can take. And there is, for each player, a payoff function ui that assigns a number to every complete combination of choices. A combination with one strategy for each player is a strategy profile, written s=(s1,s2,,sn), and the whole point of the payoff function is that it takes the entire profile as its argument: ui(s), not ui(si). There is a standard shorthand for what everybody except player i is doing, s-i, so a payoff is written ui(si,s-i) when the emphasis is on separating your own choice from the rest.

Those three ingredients together are the game in normal form, sometimes called strategic form. A two-player game with a short strategy list each is drawn as a grid: one player picks the row, the other picks the column, and the cell where they meet holds two numbers, the row player's payoff first and the column player's second. The convention on the order of those two numbers is universal and worth committing to memory, because reading a grid with them the wrong way round produces confident nonsense.

Here is one. Two airlines fly the same route, and each sets a fare of either £120 or £90 for the season. The route carries 200 passengers whatever the fares, and it costs an airline £30 to carry one. At equal fares the traffic splits evenly. If one undercuts, it takes 170 passengers and leaves 30 to its rival. Multiplying out gives four cells, in thousands of pounds:

A's fareB charges £120B charges £90
£1209.0, 9.02.7, 10.2
£9010.2, 2.76.0, 6.0

Every number there was computed, not invented. If both hold at £120 they carry 100 passengers each at a margin of £90, which is £9,000. If A alone cuts to £90, its margin falls to £60 but it carries 170, which is £10,200, while B is left with 30 passengers at a £90 margin, which is £2,700. The table is doing real work: read across A's top row and you see A's fortune swinging between £9,000 and £2,700 on a decision that belongs to B.

Example. Suppose loyalty is weaker than assumed and the undercutter takes 185 of the 200 passengers rather than 170. What do the two off-diagonal cells become?

The margins are unchanged, £90 at the high fare and £60 at the low one, so only the passenger counts move. The undercutter earns 185×60=11{,}100 and the rival earns 15×90=1{,}350. The off-diagonal cells become 11.1, 1.35 and 1.35, 11.1 in thousands. Undercutting became more attractive and being undercut became worse, while the two diagonal cells did not move at all.

Now you. Return to the original 170/30 split, but suppose the cost of carrying a passenger rises from £30 to £50. What are the four cells now?

Answer

The margins become £70 at the high fare and £40 at the low one. Both high: 100×70=7{,}000 each. Both low: 100×40=4{,}000 each. Undercutter: 170×40=6{,}800; rival: 30×70=2{,}100. So the grid reads 7.0, 7.0 across the top left, 2.1, 6.8 and 6.8, 2.1 off the diagonal, and 4.0, 4.0 at the bottom right, in thousands of pounds.

Payoffs are utilities, not money

The numbers in the grid look like money, and in the airline example they are. That is a convenience, and taking it as the definition causes trouble as soon as the analysis involves any uncertainty at all, which from the next lesson but one it always will.

What a payoff has to represent is preference over lotteries, meaning probability distributions over outcomes. That is a stronger requirement than ranking the outcomes themselves. Von Neumann and Morgenstern's contribution alongside the games was a representation theorem: if a person's preferences over lotteries satisfy four conditions, roughly that any two lotteries can be compared, that comparisons are transitive, that a preference is not reversed by mixing both sides with a third lottery, and that there are no infinitely good or infinitely bad outcomes, then there exists a function u on outcomes such that the person prefers one lottery to another exactly when it has the higher expected value of u. The payoff numbers in a game are that u.

The practical consequence is that payoffs already contain the player's attitude to risk, so the analysis never needs to add it afterwards. Take a person whose utility for a sum of money m is u(m)=m. A coin flip between nothing and £100 has expected utility 0.50+0.5100=5, and since 25=5, that gamble is worth exactly £25 to them, against an expected money value of £50. They would rather have a certain £36, worth 36=6, than the coin flip, even though the coin flip pays more on average. Put £36 and £50 in a payoff table and you would predict the wrong choice; put 6 and 5 in it and you predict the right one.

Example. The same person, with u(m)=m, is offered a coin flip between £49 and £121. What certain sum is that gamble worth to them, and how does it compare with its average payout?

Expected utility is 0.549+0.5121=0.5(7)+0.5(11)=9. The certain sum with that utility is the one whose square root is 9, so £81. The gamble pays £85 on average, so the person would trade it for £81 in cash, giving up £4 of expected money to be rid of the risk. That £4 is the risk premium, and it is what the curvature of m encodes.

Now you. Same utility function. A gamble pays £16 with probability 0.5 and £144 with probability 0.5. What is its certainty equivalent?

Answer

0.516+0.5144=0.5(4)+0.5(12)=8, and 82=64, so the gamble is worth a certain £64. Its average payout is £80, so the risk premium is £16.

There is a converse warning. Because payoffs are utilities rather than money, a player who maximises their payoff is not thereby selfish. If you would genuinely rather split a windfall than keep it, that preference belongs in your utility numbers, and a model in which everyone maximises ui is a model in which everyone pursues what they actually want. This is a real strength of the framework and also its most abused escape hatch, because any observed behaviour whatsoever can be rationalised after the fact by adjusting the payoffs. The last lesson of this course is largely about that temptation.

What can be changed without changing the game

A utility function is not unique. If u represents someone's preferences over lotteries, so does v=αu+β for any constant α>0 and any constant β, because expectations are linear: E[v]=αE[u]+β preserves every comparison. This is called invariance under positive affine transformation, and it is the exact analogue of measuring temperature in Celsius or Fahrenheit.

So the same game can be written with different numbers. A player's payoffs may be doubled, or have 50 added, without changing a single strategic conclusion. What is not allowed is applying a transformation that is not affine, such as squaring, and what is emphatically not allowed is comparing one player's payoff with another's. A cell reading 9.0, 9.0 does not mean the players are equally well off, and a cell reading 10.2, 2.7 does not mean the first player gains four times what the second suffers. Interpersonal comparison needs an extra assumption, and the bargaining lessons later in this course will have to make one explicitly.

Example. Player 1's payoffs across four cells are 0, 1, 3, 4. A colleague writes the same game with 2, 5, 11, 14. Are these the same preferences?

Test whether one constant multiplier and one constant offset carry all four across. From 0 to 2 the offset is 2 if the multiplier is anything. From 1 to 5: α(1)+2=5 gives α=3. Check the rest: 3(3)+2=11 and 3(4)+2=14, both right. It is the transformation v=3u+2 with α=3>0, so the two tables describe the same player.

Now you. Player 2's payoffs are 1, 2, 4 and a colleague writes 1, 4, 16. Same preferences?

Answer

No. Matching the first two needs α(1)+β=1 and α(2)+β=4, so α=3 and β=-2. That predicts 3(4)-2=10 for the third, not 16. The colleague squared the payoffs, which is not an affine transformation, and it changes how the player ranks lotteries: a coin flip between 1 and 16 beats a certain 4 in the second table and loses to it in the first.

Actions, strategies, and the difference

In the airline grid a strategy is just a choice, and the words "action" and "strategy" can be used interchangeably. That stops being true the moment players move in sequence and can see what has happened.

A strategy is a complete contingent plan: it specifies what the player does at every point where they might have to act, including points that will never be reached if the plan is followed. A chess strategy in this sense is not an opening preference but a full specification of a reply to every legal position, which is why the strategy sets of real games are astronomically large and why the normal form is a theoretical device rather than something anybody writes out. The definition looks pedantic and it earns its keep later: it is what allows a sequential game to be analysed with the same equilibrium concept as a simultaneous one, and it is what makes the phrase "a threat nobody would carry out" precise.

Simultaneity matters less than it sounds, too. What the normal form assumes is not that the players move at the same instant but that neither learns the other's choice before committing. Sealed bids opened a week apart are simultaneous in the only sense the model cares about.

What the model assumes

Three assumptions are doing the work, and they should be visible rather than smuggled.

Players are rational: each has preferences satisfying the utility axioms and chooses to maximise expected payoff given their beliefs. This does not claim that people are calculating machines; it is a baseline that says what a situation rewards, and the interesting research programme of the past forty years is the catalogue of situations where the baseline misses.

The structure is common knowledge: every player knows the players, strategies and payoffs, every player knows that every player knows, and so on without limit. This is much stronger than everyone happening to know it, and the infinite regress is not decorative. Several of the arguments ahead, iterated deletion in particular, use the higher levels explicitly.

And players do not communicate or bind themselves unless the game says so. A promise, a contract or a hostage changes the game rather than the reasoning about it, which is why the lesson on commitment is about redesigning the tree rather than about being trustworthy.

Where this leaves us

The airline grid is now fully specified, and specifying it has predicted precisely nothing. Both firms would rather be in the 9.0, 9.0 cell than the 6.0, 6.0 cell, and yet neither has any way of getting there by choosing well. What is missing is a rule that takes a grid and returns a prediction.

The next lesson supplies the weakest such rule, one that needs no assumption about what your rival believes, only that they will not play something that is worse for them no matter what happens. Applied to the airline grid it gives a single answer, and the answer is the cell both firms like least.

Dominance

Predicting what a player will do usually requires knowing what they expect everyone else to do, and there is one case where it does not: an option that is worse than another whatever happens.

An option no belief can justify

The previous lesson left two airlines with a grid and no prediction. Each sets a fare of £120 or £90, and in thousands of pounds the payoffs are 9.0, 9.0 when both hold high, 6.0, 6.0 when both cut, and 10.2, 2.7 in favour of whichever one undercuts alone.

Look at the grid from A's side, one column at a time. If B holds at £120, A earns 9.0 by holding and 10.2 by cutting, so cutting is better. If B cuts to £90, A earns 2.7 by holding and 6.0 by cutting, so cutting is better again. A does not need to know, guess or estimate what B will do. Whatever B does, cutting pays more.

That is strict dominance. Strategy si strictly dominates si for player i when

ui(si,s-i)>ui(si,s-i)for every s-i

meaning for every combination of choices the other players might make. A rational player never plays a strictly dominated strategy, and the argument is airtight in a way that almost nothing else in this course is: whatever beliefs the player holds, however wrong those beliefs are, the dominating strategy yields a higher expected payoff, because it yields more against every single thing the others could do. The converse holds too, in a form worth remembering: in a finite two-player game, a strategy is strictly dominated exactly when there is no belief at all about the opponent to which it is a best reply.

The grid is symmetric, so the same argument runs for B. Cutting is dominant for both, both cut, and the prediction is 6.0, 6.0. Both firms end up £3,000 worse off than in the cell they were both trying to reach, and no one made a mistake.

The prisoner's dilemma

That structure has a name, and it is the most reproduced object in the social sciences. Merrill Flood and Melvin Dresher constructed the game at the RAND Corporation in January 1950 and ran it as a hundred-round experiment between two colleagues, who cooperated in about sixty of the hundred rounds and left irritated notes about each other's play. Albert Tucker, needing to explain the payoff structure to an audience of psychologists later that year, dressed it in the story that stuck.

Two suspects are held separately. The evidence supports a minor charge with a one-year sentence. Each is offered the same deal: testify against the other and walk free, while the other serves ten years. If both testify, the deal is withdrawn from both and each serves six years. Payoffs are negative years served, so the grid is:

Prisoner 12 stays silent2 testifies
Stays silent-1, -1-10, 0
Testifies0, -10-6, -6

Testifying strictly dominates silence: 0 beats -1 in the left column, -6 beats -10 in the right. Both testify, both serve six years, and both would have served one had they held. The outcome is Pareto dominated, meaning there is another outcome both players strictly prefer, and rational individual choice walks straight past it.

The general form uses four letters. Write T for the temptation payoff to the sole defector, R for the reward when both cooperate, P for the punishment when both defect, and S for the sucker's payoff when you cooperate alone. The game is a prisoner's dilemma exactly when

T>R>P>S

and repeated-play analysis normally adds a second condition, 2R>T+S, which says that mutual cooperation beats taking turns at exploiting each other. Both conditions are checkable arithmetic, and a great many situations described as prisoner's dilemmas in print fail one of them.

Real instances are everywhere, and they are not moral parables. Two OPEC members each gain by producing above quota while the other holds back, and the cartel price collapses when both do. Two nations each gain by arming while the other disarms. A cyclist gains by doping while the peloton stays clean, and the sport ends up with a doped peloton riding at the same relative speeds. A farmer gains by drawing from the shared aquifer, and the aquifer runs dry. In each the dominant strategy is individually correct and collectively ruinous, and the standard repairs, which are contracts, regulators, monitoring, and above all repetition, all work by changing the payoffs rather than by improving anyone's reasoning.

Example. Two firms choose to cooperate or defect. Mutual cooperation pays 4 each, mutual defection pays 2 each, and a sole defector gets 6 while the cooperator gets 1. Is this a prisoner's dilemma?

Read off T=6, R=4, P=2, S=1. The ordering condition holds: 6>4>2>1. The second condition also holds, since 2R=8 exceeds T+S=7. It is a prisoner's dilemma in the full sense, and defection strictly dominates for both, since 6 beats 4 against a cooperator and 2 beats 1 against a defector.

Now you. Same structure with mutual cooperation paying 3 each, mutual defection paying 1 each, a sole defector getting 7 and the cooperator getting 0. Does this game satisfy both conditions?

Answer

T=7>R=3>P=1>S=0, so the ordering condition holds and defection is still strictly dominant. But 2R=6 is less than T+S=7, so the second condition fails. Two players who could coordinate would do better taking turns at exploitation, averaging 3.5 each, than by cooperating every round. That matters once repetition is on the table, because a cooperative arrangement is not the best thing the pair can arrange.

Iterated deletion

Dominance is worth much more than the games in which someone has a dominant strategy, because deleting a dominated strategy leaves a smaller game in which something else may now be dominated.

Take a three-by-three game. Player 1 chooses the row from U, M and D, player 2 chooses the column from L, C and R, and the cells hold the pair of payoffs:

Player 1LCR
U4, 35, 16, 2
M2, 18, 43, 6
D3, 09, 62, 8

No player has a dominant strategy here. Start with player 2, whose payoffs are the second number in each cell. Compare column C with column R: C pays 1 against U where R pays 2, C pays 4 against M where R pays 6, C pays 6 against D where R pays 8. R is better in every row, so C is strictly dominated and player 2 will not choose it. Delete the C column.

Now look again at player 1, who is left choosing among rows against L and R only. Against L, U pays 4, M pays 2 and D pays 3; against R, U pays 6, M pays 3 and D pays 2. U beats both of the others in both columns, so M and D are now strictly dominated. Note that they were not dominated in the original game: D paid 9 against C, better than U's 5. It is the deletion of C, itself justified only by player 2's rationality, that makes them dominated. Delete M and D.

Player 2 is now choosing against U alone, where L pays 3 and R pays 2. Delete R. One cell survives: (U, L), paying 4 to player 1 and 3 to player 2.

The procedure is called iterated deletion of strictly dominated strategies, and its assumptions escalate at every round. Deleting C requires only that player 2 is rational. Deleting M and D requires that player 1 is rational and knows that player 2 is. Deleting R requires that player 2 knows that player 1 knows that player 2 is rational. That is the common knowledge assumption from the previous lesson being spent, one level per round, which is why the deeper rounds of these arguments should be trusted less than the first.

One reassurance about the procedure: for strict dominance, the order of deletion does not matter. Whatever sequence of legal deletions you make, the same set of strategies survives, because a strategy strictly dominated in the full game is still strictly dominated in any subgame you reach by deleting other strategies. That is not a small property, and the next section shows how badly it fails when "strictly" is weakened.

Example. Apply iterated deletion to this game.

Player 1LCR
T4, 22, 33, 1
M5, 13, 42, 2
B2, 01, 21, 3

Player 1's row B pays 2, 1, 1 against the three columns, while T pays 4, 2, 3. T beats B everywhere, so B goes. With rows T and M left, player 2's column C pays 3 and 4, column L pays 2 and 1, column R pays 1 and 2, so C strictly dominates both and L and R go. Player 1 then compares T's 2 with M's 3 in column C and takes M. The prediction is (M, C), paying 3 to player 1 and 4 to player 2.

Now you. In the three-by-three game solved in the section above, which single deletion has to come first, and what would go wrong if player 1 tried to delete M at the very start?

Answer

Only column C is dominated in the original game, so it must go first. Deleting M at the start is illegitimate because M is not dominated there: against C it pays 8 while U pays 5, so M is better than U in that column. Player 1 can only rule M out after ruling out C, which is a claim about player 2's rationality rather than about the payoff table alone.

Weak dominance, and why it is treated with suspicion

Strategy si weakly dominates si when it is at least as good against everything and strictly better against something. That sounds like a harmless relaxation. It is not.

Player 1LR
T1, 10, 0
M1, 12, 1
B0, 02, 1

For player 1, M weakly dominates T, since it ties at L and pays 2 against 0 at R. M also weakly dominates B, since it pays 1 against 0 at L and ties at R. Both deletions are legal, and they lead to different places.

Delete T first. Player 2, choosing between L and R against rows M and B, finds that R ties with L against M and beats it against B, so L is weakly dominated and goes. What survives pays 2 to player 1 and 1 to player 2. Now start again and delete B first instead. Player 2, choosing against rows T and M, finds that L ties with R against M and beats it against T, so R goes. What survives pays 1 to player 1 and 1 to player 2. Same game, two legal orders, and player 1's payoff is 2 or 1 depending on which was chosen.

So iterated weak dominance is not a well-defined solution procedure, and results derived with it should always say which order was used. Weak dominance is still worth having as a one-shot argument about a single strategy, and one of the most useful results in the whole subject is of that kind: in a sealed-bid second-price auction, bidding your true value weakly dominates every other bid. That argument, and the auction it lives in, arrive near the end of this course.

Dominance by a mixed strategy

A strategy can fail to be dominated by any single alternative and still be dominated by a randomisation over alternatives. Since the payoffs are utilities, the expected payoff of a randomisation is a legitimate comparison, which is exactly what the utility axioms from the previous lesson were for.

Player 1LR
U4, 00, 0
M1, 01, 0
D0, 04, 0

M is not dominated by U, which collapses to 0 against R, nor by D, which collapses to 0 against L. But consider the strategy "play U with probability 0.5 and D with probability 0.5". Against L it yields 0.5(4)+0.5(0)=2; against R it yields 0.5(0)+0.5(4)=2. Two beats one in both columns, so M is strictly dominated by the mixture and can be deleted, even though a player using it would never write down a plan that includes M's guaranteed 1.

Example. Replace M's payoffs above with 3 against L and 3 against R, leaving U and D as they were. Is M dominated by any mixture of U and D?

A mixture playing U with probability p pays 4p against L and 4(1-p) against R. To dominate M it must beat 3 in both columns, so 4p>3 and 4-4p>3, meaning p>0.75 and p<0.25 at once. No such p exists, so M survives. Its guaranteed 3 is too good to be beaten by a gamble on the extremes.

Now you. Now let U pay 5 against L and 0 against R, D pay 0 against L and 5 against R, and M pay 2 against both. Is M strictly dominated by a mixture?

Answer

The mixture pays 5p against L and 5(1-p) against R, and beating 2 in both columns needs p>0.4 and p<0.6. Any p strictly between 0.4 and 0.6 works, and p=0.5 gives 2.5 against either column, so M is strictly dominated and can be deleted.

What dominance cannot do

For all its rigour, dominance answers almost nothing. Most games have no dominated strategies at all, and iterated deletion in those games removes nothing and predicts nothing. Even when it bites, it often leaves a large rectangle of survivors rather than a single cell. It is a filter, not a solution concept, and a subject that stopped here would be able to analyse the prisoner's dilemma and very little else.

The empirical record is mixed even where the theory is sharpest. First plays of a one-shot prisoner's dilemma in the laboratory produce cooperation in something like a third to a half of cases, well above the zero the dominance argument predicts, and the rate declines but does not vanish with experience. People also play strictly dominated strategies in games where the dominance is hidden behind a step of arithmetic, which is evidence about attention rather than about preferences. Both facts return in the final lesson.

What is needed is a weaker requirement that still cuts: not "better whatever they do", which is rare, but "best given what they are actually doing". Making that idea non-circular, when what they are doing depends in turn on what you are doing, is the single most important step in the subject, and it is next.

Nash equilibrium

Dominance predicts something only in the rare game where one option beats another whatever the opponent does, so the useful question is the weaker one: what is best given what the opponent is actually doing.

Best responses

Fix what everybody else is doing and the strategic problem collapses into an ordinary decision problem. If player i knows the others are playing s-i, they simply pick whichever of their own strategies pays most. That choice is the best response:

si is a best response to s-i when ui(si,s-i)ui(si,s-i) for all siSi

Note the plural is possible: two strategies can tie, so a best response is really a set. Note also that this is a conditional statement and predicts nothing on its own, because s-i is not known.

On a grid, best responses are found by scanning. Go down each column and mark the row player's largest payoff in it; that row is the row player's best response to that column. Then go along each row and mark the column player's largest payoff in it. Any cell carrying both marks is a cell in which each player is doing the best they can given the other, and marking a grid this way takes a few seconds once the habit is formed.

Here is a game to mark. Player 1 chooses the row from T, M and B, player 2 the column from L, C and R:

Player 1LCR
T3, 10, 25, 0
M2, 34, 41, 2
B1, 03, 16, 5

Reading down player 1's payoffs, the best reply to L is T with 3, to C is M with 4, and to R is B with 6. Reading across player 2's payoffs, the best reply to T is C with 2, to M is C with 4, and to B is R with 5. Two cells carry both marks: (M, C) and (B, R).

The definition

A Nash equilibrium is a strategy profile in which every player's strategy is a best response to the others' strategies. Written out, s*=(s1*,,sn*) is an equilibrium when for every player i and every alternative si,

ui(si*,s-i*)ui(si,s-i*)

The condition is entirely about unilateral deviations. It asks each player, one at a time, whether they could do better by changing their own strategy while everyone else holds theirs fixed. If nobody can, the profile is an equilibrium. Two players changing together is not a deviation the definition considers, which is precisely why the prisoner's dilemma has a bad equilibrium: both switching to cooperation is an improvement for both, and neither switching alone is.

That gives three ways of saying the same thing, and each is worth having. An equilibrium is a profile of mutual best responses. It is a profile at which no player regrets their own choice, given what the others turned out to do. And it is a self-enforcing agreement: if the players could talk beforehand but sign nothing, an equilibrium is an agreement each would keep out of self-interest once the talking stopped. A non-equilibrium agreement is one somebody breaks the moment it becomes binding on the others alone.

The concept is John Nash's, published as a two-page note in the Proceedings of the National Academy of Sciences in 1950 while he was a graduate student, and in full in the Annals of Mathematics in 1951. It generalised von Neumann's earlier solution of zero-sum games, which is the subject of a later lesson, to games where interests are not perfectly opposed, and it is the reason a subject that had a beautiful theory of poker acquired one of oligopoly, arms control and auctions. Nash shared the 1994 Nobel Memorial Prize in Economics with John Harsanyi and Reinhard Selten, both of whom appear later in this course.

Example. In the grid above, verify that (M, C) is an equilibrium and that (T, C) is not.

At (M, C), player 1 has 4 and could switch to T for 0 or B for 3, so staying is best. Player 2 has 4 and could switch to L for 3 or R for 2, so staying is best. Neither can gain alone, so it is an equilibrium. At (T, C), player 1 has 0 and could switch to M for 4. One profitable deviation is enough to disqualify a profile, so (T, C) is not an equilibrium, even though player 2 is perfectly happy there.

Now you. Is (B, R), paying 6 to player 1 and 5 to player 2, an equilibrium of the same game?

Answer

Player 1 at (B, R) has 6, against 5 from T and 1 from M, so no gain. Player 2 has 5, against 0 from L and 1 from C, so no gain. It is an equilibrium, and it happens to pay both players more than the other equilibrium (M, C) does. Two equilibria, one better for everyone, and nothing in the definition says which one occurs. That problem gets a lesson of its own shortly.

How it sits with dominance

The two ideas do not conflict, and the relations between them are worth stating exactly.

If every player has a strictly dominant strategy, that profile is a Nash equilibrium, and it is the only one. The prisoner's dilemma is the case in point: mutual defection is the unique equilibrium, and it is unique because a dominated strategy is never a best response to anything, so no profile containing one can qualify.

More usefully, iterated deletion of strictly dominated strategies never deletes an equilibrium strategy. Suppose it did, and consider the first equilibrium strategy deleted: it is strictly dominated in the reduced game, so it is beaten against every surviving profile of the others, including the equilibrium profile of the others, which is still present because this deletion was the first. That contradicts its being a best response. So every Nash equilibrium survives deletion, and if deletion leaves exactly one cell, that cell is the unique equilibrium of the game.

The converse fails, and it fails often: survival is much weaker than equilibrium. In the grid above nothing at all is strictly dominated, so deletion leaves all nine cells, while only two of them are equilibria. Dominance is a coarse filter; equilibrium is a fine one.

What equilibrium is not

More confusion attaches to this concept than to any other in the subject, and most of it comes from four claims that sound true and are not.

It is not the best outcome. The prisoner's dilemma equilibrium is worse for both players than the cell they both abandoned. Equilibrium is a stability condition, not a welfare condition, and the entire field of mechanism design exists because the two come apart.

It is not unique. The grid above has two, coordination games have three, and repeated games, as a later lesson shows, have infinitely many. A prediction that offers several answers is a weaker prediction, and pretending otherwise by quietly picking the nicest equilibrium is one of the standard ways an applied game-theoretic argument goes wrong.

It is not a claim about what people do on first meeting a game. Equilibrium requires each player's beliefs about others to be correct, and there is no reason a stranger's first play should satisfy that. What the concept describes is a rest point: a situation nothing in the players' own interests pushes them away from. Nash himself supplied a second reading in his 1950 thesis, the "mass action" interpretation, in which players are drawn from large populations, learn the average behaviour of the other side over time, and equilibrium is where that adjustment stops. That reading demands no cleverness of the players at all, and it is the honest one to have in mind when the theory is applied to markets, animals or traffic.

It is not a recommendation to any individual. "Play your equilibrium strategy" is good advice only if the others are playing theirs. Against a weak or predictable opponent, the equilibrium strategy is usually not the best reply, and a poker player who mixes to equilibrium against a novice is leaving money on the table in exchange for being unexploitable.

Example. Find all pure-strategy equilibria of this game by marking best responses.

Player 1LCR
T5, 22, 11, 3
M3, 04, 50, 1
B2, 41, 26, 2

Player 1's best replies, column by column: against L it is T with 5, against C it is M with 4, against R it is B with 6. Player 2's best replies, row by row: against T it is R with 3, against M it is C with 5, against B it is L with 4. The only cell where the two agree is (M, C), paying 4 to player 1 and 5 to player 2. One equilibrium, and note that the largest single payoff in the table, player 1's 6 at (B, R), sits in a cell that is not one.

Now you. Find all pure-strategy equilibria of this game.

Player 1LCR
T4, 51, 22, 0
M0, 13, 45, 2
B2, 30, 16, 2
Answer

Player 1's best replies are T against L (4), M against C (3), and B against R (6). Player 2's best replies are L against T (5), C against M (4), and L against B (3). Two cells agree: (T, L) paying 4 and 5, and (M, C) paying 3 and 4. Note that (B, R) is not an equilibrium despite holding player 1's best payoff in the table, because player 2 would switch to L.

When there is no equilibrium at all

Mark the best responses of this one, a tax authority choosing whether to audit and a taxpayer choosing whether to evade. Payoffs are in thousands, the taxpayer's first:

TaxpayerAuditDo not audit
Evade-100, 8050, -50
Declare0, -200, 0

Follow the best responses round. If the authority audits, the taxpayer prefers to declare, since 0 beats -100. If the taxpayer declares, the authority prefers not to audit, since 0 beats the -20 cost of a wasted audit. If the authority does not audit, the taxpayer prefers to evade, since 50 beats 0. If the taxpayer evades, the authority prefers to audit, since 80 beats -50. Four arrows, and they form a closed cycle: every cell has somebody who wants out. No cell carries both marks, so the game has no equilibrium in the strategies as defined.

This is not an artefact of the numbers. Matching pennies, in which one player wins if two coins match and the other wins if they differ, has the same cyclical structure with the smallest possible payoffs, and so does every game whose essence is that one side wants to be predicted wrongly. Penalty kicks, serve direction in tennis, bluffing, inspection, camouflage and pursuit are all of this shape, which makes the gap in the theory a serious one rather than a curiosity.

Example. Does this game have a pure equilibrium?

Player 1LR
T3, 21, 1
B2, 00, 3

Player 1's best reply to L is T (3 against 2) and to R is also T (1 against 0), so T is dominant. Player 2's best reply to T is L (2 against 1). The cell (T, L) carries both marks and is the unique pure equilibrium. The cycle of the audit game does not appear here because one player has a strategy that is good regardless.

Now you. In the audit game, suppose the fine is raised so that evading under audit costs the taxpayer 300 rather than 100. Does a pure equilibrium appear?

Answer

No. The taxpayer still prefers to declare when audited, since 0 beats -300, and still prefers to evade when not audited, since 50 beats 0. The authority's preferences are untouched. The four arrows still form the same cycle, so the game still has no pure equilibrium. Raising the penalty changes how much is at stake without changing anybody's ranking in any column, and the next lesson shows the surprising thing it does change.

What a fix has to do

The cycling game exposes something specific. A player in the audit game does not want to be predictable, and every strategy considered so far is perfectly predictable, because it names one action. The obvious repair is to let a player choose a probability distribution over their actions instead, and then to ask what the equilibrium condition means when strategies are lotteries.

That repair is not a patch on a defective theory. It restores existence completely: with randomisation allowed, every finite game has at least one equilibrium, which is Nash's theorem and the reason his name is on the concept rather than only on the definition. Working out what those probabilities have to be, and reading the strange comparative statics they imply for penalties, audits and penalty kicks, is the next lesson.

Mixed strategies

A game in which one player wants to be predicted wrongly has no equilibrium among strategies that name a single action, because naming a single action is exactly what makes a player predictable.

Strategies that are lotteries

The previous lesson left the tax authority and the taxpayer chasing each other round a cycle of best responses, with no cell where both were content. Extend what a strategy is allowed to be and the cycle closes.

A mixed strategy for player i is a probability distribution σi over that player's pure strategies. In a two-action game it is one number, the probability of the first action. The original pure strategies are still available as the degenerate mixtures that put probability 1 on one action, so nothing is lost by the extension. The set of pure strategies that a mixed strategy plays with positive probability is called its support.

Payoffs extend by expectation, which is exactly what the utility axioms of the first lesson licensed. If player 1 plays Top with probability p and player 2 plays Left with probability q, then each cell of a two-by-two grid occurs with the product of its two probabilities, and the expected payoff is the weighted sum of the four cell payoffs. Expected payoff is linear in each player's own probabilities separately, and that linearity does all the work in what follows.

The definition of equilibrium does not change at all. A profile of mixed strategies is a Nash equilibrium when no player can raise their expected payoff by switching to any other strategy, mixed or pure, holding the others fixed.

The indifference condition

Linearity has a consequence that turns equilibrium from a search into an equation.

Suppose player 1's mixed strategy is a best response to what player 2 is doing, and suppose it puts positive probability on two pure strategies, Top and Bottom. Player 1's expected payoff is then a weighted average of the expected payoff of Top and the expected payoff of Bottom. If Top were worth strictly more, shifting probability towards Top would raise the average, so the current mixture would not be a best response. Therefore:

Every pure strategy in the support of a best response yields the same expected payoff, and that payoff is at least as large as any pure strategy outside the support.

The consequence is odd on first meeting and central to everything after it. In equilibrium a mixing player is exactly indifferent between the things they are mixing over. They are not randomising because randomising is better; they are randomising because it makes no difference to them, and because it makes a great deal of difference to the other player. Which leads directly to the mechanical rule: each player's probabilities are pinned down by the other player's indifference, not by their own.

Take the audit game, with the taxpayer's payoff first and figures in thousands:

TaxpayerAuditDo not audit
Evade-100, 8050, -50
Declare0, -200, 0

Let q be the probability the authority audits. The taxpayer's expected payoff from evading is -100q+50(1-q)=50-150q, and from declaring it is 0 whatever happens. Setting them equal gives q=1/3. Now let p be the probability the taxpayer evades. The authority's expected payoff from auditing is 80p-20(1-p)=100p-20, and from not auditing it is -50p. Setting those equal gives 150p=20, so p=2/150.133.

The equilibrium is that the taxpayer evades about 13.3 per cent of the time and the authority audits a third of returns. Check it: at q=1/3 the taxpayer earns 0 from either action, so any mixture including 0.133 is a best response, and at p=2/15 the authority earns -50(2/15)=-6.67 from either action, so its mixture is a best response too. Both conditions hold, so this is an equilibrium. It is the only one: the previous lesson showed the game has no pure equilibrium, so both players must be mixing, and the two indifference equations then have exactly one solution each.

Example. A cycling team chooses whether to dope and an anti-doping agency chooses whether to test, with the team's payoff first:

TeamTestNo test
Dope-70, 4030, -30
Ride clean0, -100, 0

Find the mixed equilibrium.

Neither cell of any column is stable, so both players mix. Let q be the probability of a test. The team's payoff from doping is -70q+30(1-q)=30-100q and from riding clean it is 0, so q=0.3. Let p be the probability of doping. The agency's payoff from testing is 40p-10(1-p)=50p-10 and from not testing it is -30p, so 80p=10 and p=0.125. The team dopes an eighth of the time, the agency tests three races in ten, and each side is indifferent given the other.

Now you. Find the mixed equilibrium of this game, with player 1's payoff first.

Player 1LR
T5, 20, 4
B2, 51, 2
Answer

Let q be the probability of L. Player 1 earns 5q from T and 2q+1(1-q)=1+q from B, equal when q=0.25. Let p be the probability of T. Player 2 earns 2p+5(1-p)=5-3p from L and 4p+2(1-p)=2+2p from R, equal when 5p=3, so p=0.6. The equilibrium payoffs are 1.25 to player 1 and 3.2 to player 2, and each can be checked twice by computing it from either of that player's two actions.

The comparative static nobody expects

Now use the model for what models are for. The obvious lever a government reaches for is a harsher penalty. Suppose evasion caught in an audit costs the taxpayer 300 rather than 100, because the extra severity is a criminal record rather than money, so it costs the taxpayer without adding anything to the authority's side of the table.

Recompute. The taxpayer's indifference now reads -300q+50(1-q)=0, giving q=50/350=1/70.143. The authority's indifference is untouched, because none of its own payoffs changed, so p is still 2/15.

The evasion rate does not move. What moves is the audit rate, which falls from 33.3 per cent to 14.3 per cent. Tripling the penalty bought no reduction in evasion whatsoever; it bought a cheaper enforcement budget, and the authority's expected payoff, -50p, is unchanged at -6.67 as well. If you want less evasion, the lever that works is on the authority's side of the table: make auditing cheaper or more accurate, and p falls.

This is the general shape of mixed equilibrium comparative statics, and it is deeply counterintuitive until the indifference condition is internalised. Your own payoffs determine the other player's behaviour; the other player's payoffs determine yours. The result generalises well beyond tax: in inspection games, doping controls, fare evasion and audit-like enforcement generally, raising the punishment reduces the amount of enforcement rather than the amount of the offence, unless the enforcer's own incentives change too. The model can be wrong, and the last lesson looks at whether it is, but the mechanism is not a quirk of the numbers chosen here.

Example. In the original audit game, suppose better software cuts the cost of a wasted audit from 20 to 5. What happens to the two probabilities?

Only the authority's payoffs changed, so only the taxpayer's behaviour moves. The authority's indifference becomes 80p-5(1-p)=-50p, so 85p-5=-50p, giving 135p=5 and p=1/270.037. The taxpayer's indifference is unchanged, so the audit rate stays at q=1/3. Evasion falls from 13.3 per cent to 3.7 per cent, which is what the penalty increase failed to achieve.

Now you. Back to the original numbers, but the amount recovered by a successful audit rises from 80 to 130. What are the new equilibrium probabilities?

Answer

Again only the authority's payoffs move, so q stays at 1/3. The new indifference is 130p-20(1-p)=-50p, so 150p-20=-50p and 200p=20, giving p=0.1. Evasion falls from 13.3 per cent to 10 per cent. Anything that improves the return on enforcement reduces the offence; anything that only punishes the offender reduces enforcement.

Solving a two-by-two

The method is worth stating as a recipe, because most of the mixed equilibria anyone computes by hand are two-by-two.

Write down the opponent's payoffs. Set the opponent's expected payoff from their first action equal to their expected payoff from their second, as a function of your probability. Solve for your probability. Then swap the roles and repeat. Two linear equations, one unknown each, and no simultaneous solving is needed because the equations decouple.

Here is a version of matching pennies with unequal stakes. Player 1 wins if the coins match, and a match on heads is worth more than a match on tails. The game is zero-sum, so only player 1's payoffs are shown, and player 2's are their negatives:

Player 1headstails
Heads2-1
Tails-11

Let p be the probability player 1 plays Heads. Player 2's expected payoff from heads is -2p+1(1-p)=1-3p, and from tails it is p-(1-p)=2p-1. Equal when 5p=2, so p=0.4. Now let q be the probability player 2 plays heads. Player 1's expected payoff from Heads is 2q-(1-q)=3q-1, from Tails it is -q+(1-q)=1-2q, and these are equal when 5q=2, so q=0.4 as well. The value of the game to player 1 is 3(0.4)-1=0.2.

Look at what that says. The action with the bigger prize, Heads, is played less than half the time by the player who wants the match. Raising the reward for matching on heads makes player 1 play heads less often, because it is player 2 who reacts to that reward, by covering heads more, and player 1's own frequency is set by player 2's payoffs, which did not change until we changed them. The pattern is the same one the audit game showed, in a smaller game where it is harder to hide.

Example. Change the payoff for matching on heads from 2 to 3, leaving the rest. What is the new equilibrium, and what is the game worth to player 1?

Player 2's indifference: -3p+(1-p)=p-(1-p), so 1-4p=2p-1, giving p=1/3. Player 1's indifference: 3q-(1-q)=-q+(1-q), so 4q-1=1-2q and q=1/3. The value to player 1 is 4(1/3)-1=1/30.333. Player 1 now plays Heads only a third of the time and is better off than before, earning 0.333 rather than 0.2.

Now you. In the original version of this game (matching on heads worth 2), what is player 1's expected payoff if they stubbornly play Heads with probability 0.5 while player 2 plays the equilibrium q=0.4?

Answer

Player 2 is playing their equilibrium mixture, so player 1 is indifferent between Heads and Tails, each worth 0.2. Any mixture of them is therefore also worth 0.2, including 0.5. Deviating costs nothing, which is the indifference condition seen from the other side: an equilibrium mixture protects you against every deviation without punishing any of them.

Nash's theorem

Every finite game has at least one Nash equilibrium in mixed strategies. That is Nash's 1950 result, and it is what makes the concept usable, because a solution concept that frequently fails to apply is not a theory of anything.

The proof is a fixed-point argument and its shape is worth knowing even without the topology. Consider the map that takes any profile of mixed strategies and returns the set of best responses to it. A Nash equilibrium is exactly a profile that is in its own image, meaning a fixed point of that map. The set of mixed strategy profiles is a closed, bounded, convex set: a product of simplices. The best-response map is convex-valued, because indifference means any mixture of best responses is a best response, and it is well behaved in the technical sense the theorem requires. Kakutani's fixed point theorem then guarantees a fixed point exists. Nash's original note used exactly this; a year later he gave a shorter proof using Brouwer's theorem instead.

Two warnings come with it. The theorem promises existence, not uniqueness, and not that the equilibrium is easy to find: computing one is hard in a precise complexity-theoretic sense for large games. And it needs finiteness, or some substitute for it. Games with infinitely many strategies can fail to have equilibria without further assumptions, which matters two lessons from now when strategy sets become intervals of real numbers.

A related counting fact, due to Robert Wilson in 1971, is that almost every finite game has an odd number of equilibria. A two-by-two game with two pure equilibria therefore usually has a third one in mixed strategies, hiding between them, and finding that third one is a standard exercise in the next lesson.

What is a player actually doing?

Nobody believes a taxpayer tosses a fifteen-sided die. There are three defensible readings of a mixed equilibrium, and applied work should be explicit about which one is meant.

The first is literal randomisation, and it is exactly right in some places. Tennis players, poker players and penalty takers deliberately vary; tax authorities and customs officers really do sample at random; the routing of Allied convoys and the scheduling of security patrols are randomised by design. Where being predictable is fatal, mixing is a conscious operational choice.

The second is a population frequency. If a fresh taxpayer is drawn each year from a large population, p=0.133 can mean that 13.3 per cent of taxpayers evade with certainty and the rest never do, and the authority faces the same expected payoff. This is Nash's own mass action interpretation, and it is the natural reading in biology, where no animal randomises but a population can settle at a stable proportion of behaviours.

The third is Harsanyi's purification, from 1973, and it is the most satisfying. Suppose each player's payoffs are subject to small private fluctuations that only they observe: a taxpayer's mood, an auditor's caseload. Then almost every player has a strict pure best response given their own private shock, so nobody randomises, and yet the proportion choosing each action, as seen from outside, converges to the mixed equilibrium of the unperturbed game as the fluctuations shrink. Mixed equilibrium is then a description of an observer's ignorance rather than of a player's dice.

Where this leaves us

Existence is now settled, and the price of settling it is that uniqueness is gone for good. The audit game had exactly one equilibrium, but the three-by-three grid of the previous lesson had two, and any game whose players want to do the same thing as each other will have several.

That is not a technical nuisance. When a game has three equilibria, the theory as it stands says all three are consistent with rational play, and offers no reason to expect one rather than another. Whether anything can break the tie, and what the experiments say happens when real groups face exactly this problem, is the next lesson.

Coordination and the selection problem

When a game has one equilibrium the theory makes a prediction, and when it has three the theory says only that any of the three is consistent with rational play, which is a much weaker thing to say.

Wanting the same thing, and wanting different versions of it

The games so far have had players pulling against each other. A large and important class does the opposite: what each player most wants is to match the other, and the only question is on what.

The purest case has no conflict at all. Two people are cut off mid-call and each must decide whether to ring back or wait. Both ringing back gives an engaged tone, both waiting gives silence, and either mismatch reconnects them. There are two equilibria, "you ring, I wait" and "I ring, you wait", both players are perfectly happy with either, and yet pairs of people fail at this constantly. Nothing in the payoffs distinguishes the two equilibria, and nothing in the theory so far tells either person which to expect.

Conflict returns as soon as the players rank the matches differently. In the game traditionally called battle of the sexes, two people want to spend the evening together but disagree about where. Payoffs, Alex first:

AlexBo goes to the operaBo goes to the match
Opera3, 10, 0
Match0, 01, 3

Both cells on the diagonal are equilibria: given that Bo is at the opera, Alex does better at the opera than alone at the football, and symmetrically. Alex prefers one equilibrium and Bo the other, and the theory as it stands has nothing to say about which occurs, or about what happens if both stand firm.

The mixed equilibrium is worse than both

Wilson's odd-number result from the previous lesson says to look for a third equilibrium, and it is there.

Let p be the probability Alex goes to the opera. Bo's expected payoff from the opera is 1p+0(1-p)=p, and from the match it is 0p+3(1-p)=3-3p. Setting these equal gives 4p=3, so p=0.75. Now let q be the probability Bo goes to the opera. Alex's payoff from the opera is 3q and from the match is 1-q, equal when 4q=1, so q=0.25.

Each player goes to their own preferred venue with probability 0.75. Alex's expected payoff is 3(0.25)=0.75, and by symmetry so is Bo's. Compare that with the pure equilibria, which pay 3 and 1. The mixed equilibrium is worse for both players than either pure one, including the one each likes least.

The reason is visible in the probabilities. They meet at the opera with probability 0.75×0.25=0.1875 and at the match with the same 0.1875, so they end up together only 37.5 per cent of the time and spend 62.5 per cent of their evenings apart. This is not a pathological case. In every coordination game the mixed equilibrium sits between the pure ones and delivers less than either, which is exactly what makes it an unattractive prediction and exactly what makes it a good description of two people who genuinely do not know what the other will do.

Example. Change the numbers so that each player gets 5 at their preferred venue and 2 at the other, with 0 for missing each other. What are the mixed equilibrium probabilities and how often do they meet?

Bo's indifference: opera pays 2p, the match pays 5(1-p), so 7p=5 and p=5/70.714. By symmetry Bo goes to the opera with probability 2/70.286. They meet with probability (5/7)(2/7)+(2/7)(5/7)=20/490.408, and each earns 5(2/7)=10/71.43, again below the 2 that even the worse pure equilibrium pays.

Now you. Suppose Alex's payoffs stay 3 and 1 but Bo becomes indifferent between venues, earning 2 at either when they are together and 0 apart. What is Alex's equilibrium mixing probability now?

Answer

Alex's probability is fixed by Bo's indifference. Bo's payoff from the opera is 2p and from the match is 2(1-p), equal when p=0.5. Alex mixes evenly despite caring three times as much about the opera as about the match, because Alex's own intensity of preference never enters the equation that determines Alex's own behaviour. Bo's probability, by contrast, comes from Alex's indifference: 3q=1-q, so q=0.25.

Coordination with risk: the stag hunt

Rousseau, in the 1755 Discourse on the Origin of Inequality, described a hunting party in which each member must stay at their post to bring down a stag, and any one of them will abandon it to chase a passing hare. The game that carries his name is the most consequential of the coordination games, because the two equilibria differ in risk rather than in who gets more.

Hunter 12 hunts stag2 chases hare
Hunt stag4, 40, 3
Chase hare3, 03, 3

Both diagonal cells are equilibria. Mutual stag hunting pays 4 each and mutual hare chasing pays 3 each, so one equilibrium is better for everybody. But the hare is safe: it pays 3 whatever the other hunter does, while the stag pays 4 or nothing.

The mixed equilibrium quantifies the risk. Hunter 2 is indifferent when 4p=3p+3(1-p), that is when 4p=3, so p=0.75. Read that as a threshold on belief: hunting the stag is a best response only if you assign at least a 0.75 probability to your partner hunting the stag too. Being 70 per cent confident in your partner is not enough.

Two named criteria pull in opposite directions here. Payoff dominance says pick the equilibrium that is better for everyone, which is stag. Risk dominance, in the form Harsanyi and Selten gave it in 1988, says pick the equilibrium that is the best response to complete ignorance about the other player, meaning a fifty-fifty belief. Against a coin flip, stag pays 2 and hare pays 3, so hare is risk dominant. The formal version compares the products of what each player loses by deviating: at the stag equilibrium the product is (4-3)(4-3)=1, at the hare equilibrium it is (3-0)(3-0)=9, and the larger product wins. Hare again.

The experiments side with risk. In John Van Huyck, Raymond Battalio and Richard Beil's 1990 study, groups of fourteen to sixteen subjects each chose an effort level from 1 to 7, with the payoff determined by the minimum effort anyone chose and by their own cost of effort, a game with seven equilibria at which everyone matches. Almost a third started at the highest effort. By the tenth round nearly three-quarters of subjects were choosing effort 1, the worst equilibrium for everybody, and essentially nobody was choosing 7. One low player is enough to punish the high ones, so a single unlucky draw sends the whole group down and it never recovers. In two-player versions the efficient equilibrium was reached often, because with one partner rather than fifteen the risk is far smaller.

Example. Suppose the stag is worth 6 to each hunter rather than 4, with everything else unchanged. What is the new belief threshold, and which equilibrium is risk dominant?

Indifference now requires 6p=3, so p=0.5: a hunter needs only even odds on their partner to be worth joining. The deviation products are (6-3)(6-3)=9 for stag and (3-0)(3-0)=9 for hare, so they are exactly tied and the criterion gives no answer. Sweetening the stag by 50 per cent moved the game from clearly risk dominated to marginal.

Now you. With the original payoffs, suppose the hare is worth only 2 rather than 3, so the grid reads 4, 4 on the stag diagonal, 0, 2 and 2, 0 off it, and 2, 2 on the hare diagonal. What is the belief threshold, and which equilibrium is risk dominant?

Answer

Indifference: 4p=2, so p=0.5. The deviation products are (4-2)(4-2)=4 for stag and (2-0)(2-0)=4 for hare, another exact tie. Cheapening the safe option does the same work as enriching the risky one, which is worth knowing when the question is how to move a group to a better equilibrium: raising the reward for cooperation and lowering the value of the outside option are substitutes.

Anti-coordination: chicken

The mirror image is a game where each player wants to do the opposite of the other. Two drivers approach head on, and each may swerve or hold their line. Holding while the other swerves wins the contest, worth 1; swerving while the other holds loses it, worth -1; both swerving is a draw at 0; both holding is a collision, worth -10.

Driver 12 swerves2 holds
Swerve0, 0-1, 1
Hold1, -1-10, -10

The two pure equilibria are the asymmetric ones, in which exactly one driver swerves. Each driver prefers the equilibrium where the other yields, which puts this game close to bargaining. The mixed equilibrium is more interesting. Let p be the probability the other driver holds. Swerving pays -p; holding pays (1-p)(1)+p(-10)=1-11p. Equal when 10p=1, so p=0.1.

So in the symmetric mixed equilibrium each driver holds one time in ten, and the two of them collide with probability 0.1×0.1=0.01. Collisions are rare and they are not zero, which is the structural point: in equilibrium the disaster has positive probability, and it must, since a driver who was certain the other would swerve would never swerve themselves.

Now raise the cost of the collision from 10 to 20. Indifference becomes -p=1-21p, giving p=0.05, and the collision probability falls to 0.0025. Doubling the severity of the crash halves the aggression and quarters the accident rate, which is the same comparative static as the audit game and this time in the direction common sense expects, because here the payoff being changed belongs to both players at once.

The biological version, introduced by John Maynard Smith and George Price in 1973, is hawk-dove, and it reads the mixture as a population: 10 per cent hawks and 90 per cent doves, a stable proportion maintained by nothing more than the fact that hawks do badly in a population of hawks.

Example. In the original chicken game, what is each driver's expected payoff in the mixed equilibrium?

At p=0.1 each driver is indifferent, so the expected payoff equals the payoff of swerving, which is -p=-0.1. Both drivers do worse than the 0 they would get by both swerving every time, and they do it while colliding one time in a hundred.

Now you. With the collision cost at 20, what is the expected payoff, and how does the pair's total fare compared with the original game?

Answer

Each driver earns -p=-0.05, so the pair loses 0.1 in total against 0.2 before. Making the crash worse made the drivers better off, because it made them less aggressive. That is the logic of deterrence in one line, and its limit is that it works only while the threat of the crash is believed.

What actually selects an equilibrium

Nothing inside the model does. Every argument that picks one equilibrium out of several is an import, and it is more honest to name the import than to pretend the equilibrium concept did the work.

Focal points. Thomas Schelling, in The Strategy of Conflict in 1960, observed that people coordinate on whatever is conspicuous, and that conspicuousness has nothing to do with payoffs. Asked where they would meet a stranger in New York with no way to communicate, most of his subjects named the information booth at Grand Central Station, and almost all named noon. Later laboratory work by Judith Mehta, Chris Starmer and Robert Sugden in 1994 confirmed the effect systematically: asked to match a partner by naming heads or tails, the overwhelming majority named heads; asked for a positive number, most named 1. Salience is cultural, arbitrary and completely effective, and it is invisible to a payoff table.

Risk dominance and history. Where salience is absent, groups tend to the safe equilibrium, as the minimum-effort experiments showed, and once a group has settled somewhere, the history itself becomes the reason to stay. Which side of the road a country drives on is a pure coordination game solved permanently by precedent, and the cost of Sweden's changeover on the morning of 3 September 1967, planned for years and executed with the roads closed, is what overriding a precedent costs.

Communication. Talk that binds nobody, called cheap talk, is powerless in the prisoner's dilemma, where a player who intends to defect is happy to promise cooperation. In a coordination game it is close to decisive, because a player who announces "opera" has no incentive to lie: they want to be believed and then matched. Laboratory coordination rates rise sharply with one-way announcements, which is why the same institution, a pre-play meeting, is useless for a cartel enforcing prices and effective for two firms agreeing a technical standard.

Where this leaves us

Multiplicity is the standing weakness of equilibrium analysis, and it gets worse rather than better as the course goes on: the lesson on repeated games produces infinitely many equilibria in a game as simple as the prisoner's dilemma.

There is exactly one large class of games where the problem does not arise at all. If the players' interests are perfectly opposed, so that one player's gain is precisely the other's loss, then every equilibrium of the game gives the same payoffs, and mixing an equilibrium strategy of one with an equilibrium strategy of the other still gives an equilibrium. In those games the prediction is unique and the theory is at its strongest, which is why it was solved twenty years before Nash. That is the next lesson.

Zero-sum games and minimax

The multiplicity problem left by coordination games disappears entirely in games of pure conflict, and that is why pure conflict was solved twenty years before anything else.

Pure opposition

A two-player game is zero-sum when u1(s)+u2(s)=0 at every cell, so one player's gain is exactly the other's loss. It is constant-sum when the two payoffs add to the same constant everywhere, which is the same thing for strategic purposes: subtracting half the constant from each player's payoffs is a positive affine transformation, and the first lesson established that those change nothing.

Only one table is needed for such a game. By convention it holds the row player's payoffs, the row player is the maximiser, and the column player, whose payoffs are the negatives, is the minimiser. The phrase "player 2 wants to minimise the number in the cell" is the entire specification of their preferences.

Genuinely zero-sum situations are less common than the phrase's popularity suggests. Most trade, most negotiation and most politics contain gains from agreement, which is precisely what makes them not zero-sum, and treating them as though they were is one of the more expensive errors in public reasoning. But games and sports are built to be zero-sum, and so are the tactical layers inside larger conflicts: which side of the goal, which route for the convoy, which taxpayer to audit.

What you can guarantee yourself

Set equilibrium aside and ask a more defensive question: what can a player guarantee, whatever the opponent does?

For each of the row player's strategies, look at the worst payoff in that row. The best of those worsts is the most the row player can guarantee, called the maximin or the row player's security level. Symmetrically, for each column, look at the largest payoff in it, since that is the worst case from the minimiser's point of view, and take the smallest of those. That is the minimax, the most the column player can be forced to concede.

Two facts follow immediately. First, maximin is always less than or equal to minimax. The row player, moving in the dark, cannot guarantee more than the column player can hold them to. Second, when the two are equal, the cell where they meet is special: its payoff is simultaneously the smallest in its row and the largest in its column, so neither player can gain by moving alone. It is a saddle point, and it is a Nash equilibrium in pure strategies.

Row playerLCR
T546
M231
B314

The row minima are 4, 1 and 1, so the maximin is 4, achieved by playing T. The column maxima are 5, 4 and 6, so the minimax is 4, achieved by playing C. They agree, the game has a value of 4, and (T, C) is a saddle point: the row player cannot do better than 4 against C, and the column player cannot do better than conceding 4 against T. Neither player needs to conceal anything. The row player could announce T in advance and lose nothing, which is the signature of a game solved in pure strategies.

Example. Find the maximin, the minimax and any saddle point of this zero-sum game.

Row playerLC
T3-1
B-22

The row minima are -1 and -2, so the maximin is -1, from playing T. The column maxima are 3 and 2, so the minimax is 2, from playing C. Since -1<2 there is no saddle point, and the gap of 3 between what the row player can guarantee and what the column player can be held to is unresolved. No cell is simultaneously a row minimum and a column maximum, and following best responses round the four cells produces a cycle.

Now you. Find the maximin and minimax of this game, and say whether it has a saddle point.

Row playerLCR
T623
B145
Answer

Row minima are 2 and 1, so the maximin is 2 from T. Column maxima are 6, 4 and 5, so the minimax is 4 from C. They differ, so there is no saddle point and the game is not solvable in pure strategies. The value, once mixing is allowed, must lie between 2 and 4.

The minimax theorem

The gap between what a player can guarantee and what they can be held to is a gap about predictability, so mixing closes it. That is von Neumann's theorem, proved in the 1928 paper "Zur Theorie der Gesellschaftsspiele" and the foundation of the whole subject:

In any finite two-player zero-sum game, the maximin over mixed strategies equals the minimax over mixed strategies. The common number is the value of the game, and each player has a mixed strategy that guarantees it.

The maximiser has a mixture that yields at least v against every column, and the minimiser has one that concedes at most v against every row. Von Neumann considered this the founding theorem of game theory, and remarked that as far as he could see there could be no theory of games without it.

Computing the value in a two-by-two game with no saddle point uses the indifference method from the earlier lesson, remembering that in a zero-sum game one table serves both players. Take the game above with entries 3 and -1 in the top row, -2 and 2 in the bottom.

Let p be the probability of T. The column player's payoff from L is -3p+2(1-p)=2-5p, and from C it is p-2(1-p)=3p-2. Equal when 8p=4, so p=0.5. Let q be the probability of L. The row player's payoff from T is 3q-(1-q)=4q-1, and from B it is -2q+2(1-q)=2-4q. Equal when 8q=3, so q=0.375. The value is 4(0.375)-1=0.5.

That number is the point of the exercise. Playing T and B evenly guarantees the row player an expected 0.5 whatever the column player does, up from the -1 they could guarantee without mixing, and the column player playing L three-eighths of the time holds them to exactly that, down from the 2 they would otherwise concede. The 3-wide gap has closed to a single number.

Example. Solve this zero-sum game: 4 and -2 in the top row, -1 and 1 in the bottom.

Let p be the probability of T. The column player gets -4p+(1-p)=1-5p from L and 2p-(1-p)=3p-1 from C, equal when 8p=2, so p=0.25. Let q be the probability of L. The row player gets 4q-2(1-q)=6q-2 from T and -q+(1-q)=1-2q from B, equal when 8q=3, so q=0.375. The value is 6(0.375)-2=0.25. Check it the other way: 1-2(0.375)=0.25, as it must be.

Now you. Solve the game with 2 and 5 in the top row, 6 and 1 in the bottom.

Answer

Let p be the probability of T. The column player gets -2p-6(1-p)=4p-6 from L and -5p-(1-p)=-4p-1 from C, equal when 8p=5, so p=0.625. Let q be the probability of L: the row player gets 2q+5(1-q)=5-3q from T and 6q+(1-q)=1+5q from B, equal when 8q=4, so q=0.5. The value is 5-3(0.5)=3.5, and the check from the other row gives 1+5(0.5)=3.5.

Why the prediction is unique

Zero-sum games have three properties that no other class has, and together they are what make the theory unambiguous.

Every equilibrium has the same value. If (σ1,σ2) and (τ1,τ2) are both equilibria, both pay the row player exactly v, because both must simultaneously guarantee at least v and concede at most v. There is no equivalent of the battle of the sexes, where one equilibrium pays 3 and another pays 1.

Equilibrium strategies are interchangeable. If those two profiles are equilibria, so are the crossed pairs (σ1,τ2) and (τ1,σ2). A player therefore does not need to know which equilibrium the opponent has in mind, because any equilibrium strategy of theirs works against any equilibrium strategy of yours. The coordination problem that dominated the previous lesson simply cannot arise.

And the equilibrium strategy is exactly the security strategy. In general games these are different things: the maximin strategy is paranoid and the equilibrium strategy is best-responding, and they usually disagree. In zero-sum games they coincide, so playing for equilibrium and playing safe are the same act. This is why "play the equilibrium strategy" is genuinely good advice in poker or in a penalty shootout, and only conditionally good advice anywhere else.

One more connection is worth naming. Finding the value of a zero-sum game is a linear program, and the two players' problems are dual to each other, with the minimax theorem corresponding exactly to the duality theorem. The link was noticed by von Neumann in conversation with George Dantzig in 1947, and it means that any zero-sum game, however large, can be solved by a standard algorithm rather than by cleverness.

Do professionals actually play minimax?

The theory is sharp enough to test, and it has been. The best-known test is Ignacio Palacios-Huerta's 2003 study of 1,417 penalty kicks from professional league play in Spain, Italy and England between 1995 and 2000. A penalty is close to a genuine zero-sum game: the kicker wants to score, the goalkeeper wants a save, and both commit before seeing the other, since a shot struck at 100 km/h crosses the 11 metres in about four tenths of a second.

Grouping directions into two sides and estimating the scoring probability in each of the four combinations gives this table, as percentages, with the kicker as the maximiser:

KickerKeeper goes leftKeeper goes right
Left58.3094.97
Right92.9169.92

There is no saddle point: whichever way the keeper goes, the kicker wants to go the other way, and the keeper wants to follow. Solve it. The keeper is indifferent when 58.30p+92.91(1-p)=94.97p+69.92(1-p), where p is the probability the kicker goes left. That gives 22.99=59.66p, so p=0.385. The kicker is indifferent when 58.30q+94.97(1-q)=92.91q+69.92(1-q), giving 25.05=59.66q, so q=0.420. The value of the game is a scoring probability of 79.6 per cent.

The frequencies actually observed in the data were close to 0.40 for the kicker and 0.42 for the keeper, within a couple of percentage points of the predicted 0.385 and 0.420. That is a strong result, and it is not the strongest part of the study. Equilibrium also requires the choices to be serially independent, meaning that a kicker who went left last time is no more or less likely to go left this time, and this is where amateurs reliably fail: people asked to produce random sequences alternate too much. In the professional data the sequences pass that test too. Barry O'Neill had found the same in a 1987 laboratory card game, and Mark Walker and John Wooders found it in the serve directions of top tennis players in 2001.

Example. In the penalty table, suppose better goalkeeper coaching raises the save rate when the keeper dives left against a left-footed shot, so that the top-left entry falls from 58.30 to 50.00. What happens to the kicker's equilibrium probability of going left?

The kicker's probability is fixed by the keeper's indifference, which now reads 50.00p+92.91(1-p)=94.97p+69.92(1-p). That gives 92.91-42.91p=69.92+25.05p, so 22.99=67.96p and p=0.338. Improving the keeper's left-side save rate makes the kicker go left less often, from 38.5 per cent to 33.8 per cent, exactly the pattern from the audit game: your own payoffs move the other player.

Now you. With the original table, what would the kicker's expected scoring rate be if they went left every time against a keeper still playing q=0.420?

Answer

58.30(0.420)+94.97(0.580)=24.49+55.08=79.6 per cent, the value of the game. Against an opponent playing their equilibrium mixture every strategy in the support yields the same thing, so a kicker who abandons mixing loses nothing immediately. What they lose is protection: a keeper who noticed the pattern would move to diving left always, and the scoring rate would collapse to 58.3 per cent.

The limits of a beautiful result

Everything in this lesson depends on the interests being exactly opposed, and that assumption is fragile in a specific way: it is not preserved by adding anything the players care about jointly. Two firms in a price war are close to zero-sum over market share and not at all zero-sum over whether the war happens. A penalty kick is zero-sum; the match around it, in which a draw may suit both teams, is not.

The theorem also says nothing about how a value is divided when there is a surplus to divide, because in a zero-sum game there is no surplus by construction. Everything from here on is about games where there is one: firms choosing quantities, players moving in sequence, partners deciding whether to cooperate, and bargainers splitting a pie.

The immediate next step is smaller and more concrete. Every game so far has had a strategy set that could be listed as rows and columns, and real economic choices are numbers on a continuum: a price, a quantity, a bid, a level of effort. Best responses then become functions rather than table entries, and the equilibrium is found by solving equations. That is the next lesson.

Competing on quantity and price

Every game so far has had a strategy set short enough to write as rows and columns, and the choices that matter most in economics are numbers on a continuum: a price, a quantity, a bid, an amount of effort.

When the grid becomes a line

Nothing in the definition of equilibrium requires a finite strategy set. A best response is still the strategy maximising your payoff given the others', and an equilibrium is still a profile of mutual best responses. What changes is the arithmetic: with a continuum of strategies, the best response is found by maximising a function rather than by scanning a column, and it comes out as a best-response function giving your optimal choice for each possible choice of the other player. The equilibrium is then the solution of a system of equations, and it can often be written in closed form.

One warning carries over from Nash's theorem, which required finiteness. With infinite strategy sets, existence needs the payoff functions to be reasonably behaved, and the models here satisfy that: strategies range over a closed interval and payoffs are smooth and concave in one's own choice. When those conditions fail, equilibria can genuinely fail to exist, and the Bertrand model at the end of this lesson shows a payoff function with a jump in it doing something violent.

Cournot's duopoly

Antoine Augustin Cournot published the first equilibrium analysis in economics in 1838, more than a century before Nash, using two owners of mineral springs who each decide how much water to pump. The model still carries his name and is still the standard workhorse for quantity competition.

Two firms produce an identical good. Firm 1 chooses quantity q1, firm 2 chooses q2, and the market price is set by total supply through the inverse demand curve

p=a-b(q1+q2)

Each unit costs c to produce, and the firms choose simultaneously. Take concrete numbers to keep the algebra honest: a=100, b=1 and c=40, with quantities in thousands of units and prices in pounds.

Firm 1's profit is revenue minus cost:

π1=(p-c)q1=(100-q1-q2-40)q1=(60-q2-q1)q1

Hold q2 fixed and read that as a function of q1 alone. It is a downward parabola through the origin, zero at q1=0 and again at q1=60-q2, so its peak sits halfway between those roots. That gives the best-response function without a single derivative:

q1=60-q22

The interpretation is worth pausing on. If firm 2 produces nothing, firm 1 produces 30, which is exactly the monopoly quantity: a monopolist maximises (60-Q)Q, another parabola with roots 0 and 60, peaking at Q=30 and yielding a price of £70 and a profit of £900. Every unit firm 2 adds pushes firm 1 back by half a unit. Quantities are strategic substitutes: more of yours makes me want less of mine.

Firm 2's best response is the mirror image, q2=(60-q1)/2. In equilibrium both hold at once, and by symmetry q1=q2=q:

q=60-q23q=60q=20

So each firm produces 20, total output is 40, the price is 100-40=£60, and each firm earns (60-40)(20)=£400.

Line the three market structures up. A monopolist produces 30 at £70 and earns £900. The Cournot duopolists produce 40 between them at £60 and earn £800 between them. Perfect competition would drive the price to marginal cost, £40, with output 60 and no profit at all. Duopoly sits between monopoly and competition on every measure, which is the result Cournot was after and the reason the model survived.

Example. With the same demand and costs, what is firm 1's best response if firm 2 produces 30, and what does firm 1 earn?

The best-response function gives q1=(60-30)/2=15. Total output is 45, so the price is 100-45=£55, and firm 1 earns (55-40)(15)=£225. Firm 2, which flooded the market, earns (55-40)(30)=£450. Producing more than the equilibrium quantity is profitable for firm 2 here precisely because firm 1 is expected to accommodate it, which is a hint about what happens when one firm can move first.

Now you. Demand is p=120-2Q and each unit costs £30. What is the Cournot equilibrium quantity per firm, the price, and each firm's profit?

Answer

Firm 1's profit is (120-2q1-2q2-30)q1, a parabola with roots at q1=0 and q1=(90-2q2)/2, so the best response is q1=(90-2q2)/4. Setting q1=q2=q gives 4q=90-2q, so q=15. Total output is 30, the price is 120-60=£60, and each firm earns (60-30)(15)=£450.

The cartel and why it does not hold

The two firms would both prefer the monopoly outcome. Splitting the monopoly quantity of 30 gives each 15 units at a price of £70, so each earns (70-40)(15)=£450, against the £400 of the Cournot equilibrium. Every pound of that gain is available for the taking, and yet it is not an equilibrium.

Check it with the best-response function. If firm 2 sticks at 15, firm 1's best response is (60-15)/2=22.5. Total output becomes 37.5, the price falls to £62.50, and firm 1 earns (62.5-40)(22.5)=£506.25 while firm 2 earns (62.5-40)(15)=£337.50. Firm 1 gains £56.25 by cheating and firm 2 loses £112.50.

That is a prisoner's dilemma with a continuum of strategies. Both firms prefer the cartel to the equilibrium, each gains by defecting from the cartel while the other holds, and the unique equilibrium of the one-shot game is the outcome both like less. The numbers £450, £506.25 and £400 are the R, T and P of the second lesson, computed rather than assumed, and they return in the lesson on repetition, where they determine exactly how patient a cartel has to be to survive.

Example. Suppose the two firms agree to hold at 15 each and firm 1 cheats by producing 25 rather than its best response of 22.5. Does firm 1 still gain, and by how much?

Total output is 40, so the price is £60 and firm 1 earns (60-40)(25)=£500, which beats the £450 of holding but falls short of the £506.25 available from cheating optimally. The interesting part is firm 2, which now earns (60-40)(15)=£300, worse than the £337.50 it lost under optimal cheating. Overshooting hurts the cheat a little and the victim a lot.

Now you. With the same demand and costs, at what quantity for firm 1 would the cartel-breaking profit fall back to the £450 the cartel promised, given firm 2 holds at 15?

Answer

Firm 1's profit against q2=15 is (100-q1-15-40)q1=(45-q1)q1, and setting that equal to 450 gives q12-45q1+450=0. The roots are q1=(45±2025-1800)/2=(45±15)/2, so q1=15 or q1=30. Cheating pays for every quantity strictly between 15 and 30, and beyond 30 the firm has flooded its own market so badly that it would have done better keeping the agreement.

More firms, less market power

Nothing restricts the model to two firms. With n identical firms the same argument runs: firm i's profit is (a-c-bQ-i-bqi)qi, a parabola in qi with roots 0 and (a-c-bQ-i)/b, peaking halfway. Imposing symmetry, so that each of the other n-1 firms produces the same q, gives

q=a-c(n+1)bandp=a+ncn+1

With the running numbers, q=60/(n+1) and p=100-60n/(n+1). One firm produces 30 at £70. Two produce 20 each at £60. Three produce 15 each at £55 and earn £225 each. Five produce 10 each at £50 and earn £100 each. Ten produce about 5.45 each at £45.45, earning under £30 each. As n grows the price approaches marginal cost from above, so perfect competition is not a separate model but the limit of this one.

The markup over cost is 60/(n+1), so most of it goes early: the second firm removes £10 of the monopolist's £30, the third another £5, while doubling from ten firms to twenty removes only £2.60. Competition policy in concentrated industries is fighting for the steep part of that curve, which is also where mergers do their damage.

Bertrand, and the paradox

In 1883 Joseph Bertrand reviewed Cournot's book and objected that firms do not choose quantities, they choose prices, and that changing the strategy variable changes the answer completely.

Two firms sell an identical good, both with marginal cost c=40, and each posts a price. Consumers buy from the cheaper, splitting evenly on a tie. What is the equilibrium?

Suppose firm 2 posts £60. Firm 1 posting £60 splits the market and earns half the profit; firm 1 posting £59.99 takes all of it and earns nearly twice as much. So £60 is not an equilibrium, and neither is any price above £40 by the same argument. Nor is any price below £40, which loses money on every unit. The only equilibrium is both firms pricing at exactly £40, marginal cost, earning nothing at all, while consumers buy 60 units at the competitive price.

This is the Bertrand paradox: two firms are enough to produce the perfectly competitive outcome. It is a paradox because real duopolies plainly do not price at marginal cost, so one of the assumptions must be doing damage. Three of them are, and each repair is a real industry.

Capacity. The undercutting argument assumes the winner can serve the entire market. If each firm can supply only 35 units of a 60-unit market, undercutting no longer captures everything, and the residual demand left to the loser supports a price above cost. David Kreps and José Scheinkman showed in 1983 that firms choosing capacity first and prices afterwards reproduce the Cournot outcome exactly, which reconciles the two models: Cournot describes competition when capacity is the real decision, Bertrand when it is not.

Differentiation. If the goods are not identical, a small price cut does not steal the whole market. Suppose firm i faces demand qi=100-2pi+pj, so its sales fall with its own price and rise with its rival's. Its profit is (pi-40)(100-2pi+pj), a parabola in pi with roots at pi=40 and pi=(100+pj)/2, peaking halfway, so the best response is pi=(180+pj)/4. Symmetry gives 4p=180+p, so p=£60, each firm sells 40 units and earns £800. Prices are strategic complements here, in contrast to Cournot quantities: a rival's price rise makes you want to raise yours. Nothing about the goods changed except that they stopped being interchangeable, and the margin came back.

Repetition. Undercutting is profitable today and provokes a price war tomorrow, and the lesson on repeated games shows precisely how patient firms have to be for the threat of that war to sustain a price above cost.

Example. In the differentiated model above, what is firm 1's best response if firm 2 prices at £50, and what does firm 1 sell?

p1=(180+50)/4=£57.50. Its sales are 100-2(57.5)+50=35 units, and its profit is (57.5-40)(35)=£612.50. Note that firm 1 does not match the £50: with differentiated goods, being dearer than your rival is compatible with selling a lot.

Now you. In the same differentiated model, suppose firm 2 prices at £80. What is firm 1's best response, and does firm 1 want to undercut all the way?

Answer

p1=(180+80)/4=£65, selling 100-130+80=50 units for a profit of (65-40)(50)=£1{,}250. Firm 1 raises its own price when firm 2 does, rather than undercutting, because the extra margin outweighs the sales lost to a differentiated rival. That is what makes prices strategic complements, and it is the reason price rises spread through a differentiated industry in a way quantity increases do not.

What the models assume

Both models assume the two firms move at the same time and know each other's costs and the demand curve exactly, and both are one-shot. Each of those assumptions is doing real work, and each is repaired in a later lesson: sequence in the next one, repetition three lessons on, and private information near the end.

Two further limits deserve naming now. Constant marginal cost is a strong assumption, and rising marginal cost changes the best-response slopes without changing the method. And the equilibrium concept still delivers a single prediction here only because these particular payoff functions are concave with a unique intersection of best responses. Add capacity limits, fixed costs or discrete plant sizes and multiple equilibria come back.

The most interesting assumption to drop is simultaneity. The example above showed firm 1 accommodating a rival who produced 30, which suggests that a firm able to commit to a large output first might do very well out of that accommodation. Working out whether it can, and by how much, requires a way of writing down who moves when. That is the next lesson.

Moving in sequence

Every game so far has assumed that neither player learns anything before choosing, and the moment one player can see what the other did, the analysis and the answer both change.

The extensive form

A game in which moves happen in a known order is written as a tree rather than a grid. The tree starts at a single node, each node is labelled with the player who moves there, each branch leaving it is an action, and every path through the tree ends at a terminal node carrying one payoff for each player. This is the extensive form.

Take the game that runs through this lesson and the next. An entrant decides whether to enter a market. If it stays out, the entrant earns 0 and the incumbent keeps its monopoly, worth 10. If it enters, the incumbent decides whether to fight a price war or accommodate. Fighting leaves the entrant with -2 and the incumbent with 2. Accommodating gives the entrant 2 and the incumbent 5. Two decision nodes, three terminal nodes, six numbers, and the whole strategic situation is specified.

Simultaneous moves fit in the same picture using an information set: a group of nodes the player on move cannot tell apart. If the incumbent must choose without knowing whether entry has occurred, its two nodes are enclosed in one information set and it must choose the same action at both, which is exactly the normal form again. A game in which every information set is a single node is a game of perfect information, and those are the games this lesson solves.

A strategy is a complete plan

The definition of a strategy from the first lesson now earns its keep. A strategy for a player is a complete contingent plan: an action at every information set where that player might have to move, including ones that will not be reached.

In the entry game the entrant has two strategies, In and Out. The incumbent has one decision node and so also has two, Fight and Accommodate. But suppose the incumbent could react differently to two different kinds of entry, say small and large: it would then have two decision nodes and four strategies, one for each combination of replies. In general a player with k decision nodes and two actions at each has 2k strategies, which is why the normal form of a chess-like game is a theoretical object rather than a table anyone writes.

The insistence on specifying an action at unreached nodes looks like bookkeeping and is the opposite. It is what makes the sentence "the incumbent would have fought if the entrant had come in" a statement with a truth value, and the whole of the next lesson turns on whether such statements are believable.

Writing the entry game as a grid gives:

EntrantIncumbent fightsIncumbent accommodates
In-2, 22, 5
Out0, 100, 10

The row Out is flat, because when the entrant stays out the incumbent's plan never gets used. Mark the best responses. Against Fight the entrant prefers Out; against Accommodate it prefers In; against In the incumbent prefers Accommodate; against Out it is indifferent, so both of its strategies are best responses. Two cells carry both marks: (In, Accommodate) and (Out, Fight). The grid says both are Nash equilibria. The tree, as the next section shows, says only one of them can happen.

Example. A tree has player 1 choosing A or B at the first node. After A, player 2 chooses between C and D; after B, player 2 chooses between E and F. How many strategies does each player have, and how large is the normal form?

Player 1 has one decision node with two actions, so two strategies. Player 2 has two decision nodes with two actions at each, so four strategies: (C if A, E if B), (C if A, F if B), (D if A, E if B) and (D if A, F if B). The normal form is a two-by-four grid with eight cells, describing a tree that has only four terminal nodes.

Now you. Player 1 moves first with three actions. After each, player 2 chooses between two actions. How many strategies does player 2 have?

Answer

Player 2 has three decision nodes with two actions at each, so 23=8 strategies. The count multiplies across nodes rather than adding, which is why strategy sets in extensive games explode: a player with ten binary decision nodes has 1,024 strategies.

Backward induction

A game of perfect information is solved from the leaves inward. Find a decision node all of whose branches lead straight to terminal nodes, and settle what the player there would do, which is a simple comparison of numbers. Replace that node with the payoffs it produces. The tree is now shorter by one layer. Repeat.

In the entry game there is one node to settle: the incumbent, having been entered upon, compares 2 from fighting with 5 from accommodating, and accommodates. Replace that subtree with the pair (2, 5). Now the entrant chooses between 0 from staying out and 2 from entering, and enters. The prediction is entry followed by accommodation, paying 2 and 5.

This is backward induction, and it is the single most-used technique in applied game theory. Its logic is that a plan for the future has to be one you would actually carry out when the future arrives, which is why the (Out, Fight) equilibrium of the grid does not survive: it requires the incumbent to fight, and if entry ever happened, it would not.

Two classical results underwrite the method. Ernst Zermelo proved in 1913 that in chess either White can force a win, or Black can, or both can force at least a draw, entirely by this kind of argument and without knowing which. Harold Kuhn generalised it in 1953: every finite game of perfect information has a Nash equilibrium in pure strategies, and backward induction finds one. No mixing is needed, which is a striking contrast with the simultaneous games of the earlier lessons, and the reason is that a player who moves second never faces uncertainty about what came before.

Example. Player 1 chooses A or B. After A, player 2 chooses C, paying (3, 1), or D, paying (1, 2). After B, player 2 chooses E, paying (2, 4), or F, paying (0, 0). Solve it.

Work from the leaves. At the node after A, player 2 compares 1 from C with 2 from D and chooses D, so A is worth (1, 2). At the node after B, player 2 compares 4 from E with 0 from F and chooses E, so B is worth (2, 4). Player 1 now compares 1 from A with 2 from B and chooses B. The outcome is (2, 4), and player 2's full strategy is "D if A, E if B", which specifies a choice at the node that is never reached.

Now you. Player 1 chooses L or R. After L, player 2 picks between (5, 1) and (2, 3). After R, player 2 picks between (3, 2) and (4, 4). Solve it, and say what player 1 would get if player 2 could somehow commit in advance to the other action after L.

Answer

After L, player 2 takes 3 rather than 1, so L is worth (2, 3). After R, player 2 takes 4 rather than 2, so R is worth (4, 4). Player 1 chooses R and the outcome is (4, 4). If player 2 could commit to the (5, 1) branch after L, player 1 would prefer L, earning 5 rather than 4, and player 2 would earn 1 rather than 4. So player 2 has no interest in making that commitment here, but a player who could commit to a worse-for-themselves action can sometimes make the other side move their way. That is the subject of the next lesson.

The value of moving first

Return to the two firms of the previous lesson: demand p=100-Q, unit cost £40, and simultaneous quantity choices giving 20 each, a price of £60 and £400 each. Now let firm 1 choose first, publicly and irreversibly, before firm 2 chooses. This is Heinrich von Stackelberg's 1934 model, and it is solved by backward induction.

Start at the end. Firm 2 sees q1 and plays its best response, which the previous lesson derived: q2=(60-q1)/2. Firm 1 knows this, so it does not treat q2 as fixed; it substitutes the reaction into its own profit:

π1=(60-q1-60-q12)q1=(60-q12)q1

That is a parabola in q1 with roots at 0 and 60, so it peaks at q1=30. Firm 2 then produces (60-30)/2=15, total output is 45, the price is £55, and profits are (55-40)(30)=£450 for the leader and (55-40)(15)=£225 for the follower.

Compare the two regimes. The leader earns £450 against the £400 it would get moving simultaneously, so there is a genuine first-mover advantage of £50. The follower earns £225, well short of £400. Total output rises from 40 to 45 and the price falls from £60 to £55, so consumers gain, while total industry profit falls from £800 to £675.

Why does moving first help? Not because of information, since the follower learns something the leader does not. It helps because moving first is a commitment. Firm 1 produces 30, which is not a best response to 15, and it gets away with it because 15 is already fixed by the time firm 2 responds. In the simultaneous game, threatening to produce 30 would be empty: firm 2 would produce 15 anyway and firm 1 would then wish it had produced 22.5. Sequence converts a wish into a fact.

Example. Demand is p=120-2Q and unit cost is £30. Solve the Stackelberg game and compare the leader's profit with the £450 each firm earns in the simultaneous equilibrium.

The follower's best response, from the previous lesson, is q2=(45-q1)/2. The leader's profit is (90-2q1-2q2)q1, and substituting gives (90-2q1-45+q1)q1=(45-q1)q1, a parabola peaking at q1=22.5. Then q2=11.25, total output is 33.75, the price is 120-67.5=£52.50, and profits are (52.5-30)(22.5)=£506.25 for the leader and £253.13 for the follower. The leader gains £56.25 over the simultaneous outcome and the follower loses £196.88.

Now you. In the original Stackelberg game, what would the leader earn if it ignored the reaction function and simply produced the Cournot quantity of 20?

Answer

The follower would respond with (60-20)/2=20, total output would be 40, the price would be £60 and the leader would earn £400, exactly the simultaneous-game profit. Moving first is worth nothing unless the leader exploits it, and exploiting it means deliberately producing more than the quantity that would be a best response after the fact.

Moving first is not always better

The Stackelberg result is often quoted as though first movers always win, and the previous lesson supplies the counterexample. The advantage comes from committing to something that pushes the follower's best response in your favour, and whether that is possible depends on whether the strategic variables are substitutes or complements.

With quantities, more of yours makes the rival want less, so committing to a lot is aggressive and profitable. With differentiated prices it is the other way round: a rival's higher price makes you want a higher price too. A price leader in the differentiated model of the previous lesson, where each firm faces qi=100-2pi+pj with cost £40, maximises (p1-40)(100-2p1+(180+p1)/4) and sets £61.43, which is above the simultaneous £60. The follower then prices at £60.36 and earns £828.83, while the leader earns £803.57. Both do better than the £800 of the simultaneous game, and the follower does best of all. Announcing a high price first is an invitation the rival accepts while slightly undercutting.

So the useful statement is not "move first" but "commit to something the other side must accommodate", and whether such a commitment exists is a property of the game rather than of the clock.

Where backward induction becomes uncomfortable

The method is only as good as the assumption that every player will apply it at every future node, and there is a small game designed to make that assumption look ridiculous.

In Robert Rosenthal's centipede game, two players alternate. A pot starts at 50 pence, split 40 to the player on move and 10 to the other. Passing doubles the pot and hands the move over. In the four-move version used by Richard McKelvey and Thomas Palfrey in 1992, taking at the first node gives (0.40, 0.10), at the second (0.20, 0.80), at the third (1.60, 0.40) and at the fourth (0.80, 3.20), with (3.20, 0.80) if the last player passes.

Backward induction is brutal. At the last node player 2 takes 3.20 rather than passing for 0.80. Knowing that, at the third node player 1 takes 1.60 rather than the 0.80 that passing would bring. At the second node player 2 takes 0.80 rather than 0.40. So at the first node player 1 takes 0.40 rather than the 0.20 that passing would bring, and the game ends immediately with 50 pence on the table when four pounds was available.

Almost nobody plays this way. In McKelvey and Palfrey's experiments only a small minority took at the first opportunity, most games ran several nodes, and the pot regularly grew. The theory's own logic explains the discomfort: player 2's move at the second node is only reached if player 1 has already done something backward induction says is irrational, and player 2 then has to decide what to believe about a player who has just violated the assumption the whole argument rests on. Backward induction has no answer to that, and the final lesson of this course returns to what the experiments show and what the models built to fit them look like.

Where this leaves us

The tree solved the entry game and threw away one of its two Nash equilibria. The one it threw away, in which the incumbent's plan to fight keeps the entrant out, is exactly the kind of arrangement that real firms, unions, governments and parents rely on: a threat that would be costly to carry out but works if believed.

Making the distinction precise, which means saying which Nash equilibria of the grid survive when the tree is taken seriously, is the concept of subgame perfection. And once threats are classified into credible and empty, the natural follow-up is how a player might make an empty threat into a real one by destroying their own options. That is the next lesson.

Credible threats and commitment

The entry game had two Nash equilibria on the grid and only one survived the tree, and the difference between them is the difference between a threat that would be carried out and a threat that would not.

The equilibrium the tree rejects

Recall the game. An entrant chooses In or Out; if it enters, the incumbent chooses Fight or Accommodate. Staying out pays the entrant 0 and the incumbent 10. Fighting pays -2 and 2. Accommodating pays 2 and 5.

EntrantIncumbent fightsIncumbent accommodates
In-2, 22, 5
Out0, 100, 10

Both (In, Accommodate) and (Out, Fight) are Nash equilibria. Check the second one against the definition: given that the incumbent's plan is to fight, the entrant's best reply is to stay out, since 0 beats -2. Given that the entrant stays out, the incumbent's payoff is 10 whatever it plans to do, so planning to fight is a best response too. Nothing in the definition of Nash equilibrium is violated.

And yet the plan is nonsense. If the entrant did come in, the incumbent would compare 2 with 5 and accommodate. The threat works only for as long as it is never tested, and the reason Nash equilibrium tolerates it is that the definition only asks whether a player could gain by deviating given the others' strategies, and the entrant's staying out means the incumbent's promise is never priced.

Subgame perfection

Reinhard Selten's repair, introduced in 1965, is to require the strategies to be an equilibrium not just in the game as a whole but in every part of it that could be reached.

A subgame is any node of the tree together with everything that follows it, provided no information set is cut in half by taking it. The whole game is a subgame of itself, and in a game of perfect information every node starts one. A subgame perfect equilibrium is a strategy profile that induces a Nash equilibrium in every subgame.

The entry game has two subgames: the whole thing, and the small one that begins after entry in which the incumbent alone chooses. In that small subgame, the only equilibrium is Accommodate, because the incumbent is the only player and 5 beats 2. So (Out, Fight) is not subgame perfect: it fails in the subgame after entry. Only (In, Accommodate) survives, and it is exactly what backward induction gave.

For finite games of perfect information the two ideas coincide: backward induction is subgame perfection carried out mechanically. The general definition earns its keep when information sets are not all singletons, because then the tree contains simultaneous-move subgames that must be solved with the earlier machinery and cannot be settled by comparing two numbers.

The concept refines rather than replaces. Every subgame perfect equilibrium is a Nash equilibrium, and some Nash equilibria are not subgame perfect. What is discarded is precisely the set of profiles held up by threats and promises that would not be honoured.

Example. A supplier and a buyer. The buyer chooses whether to invest £4 in equipment that only this supplier's parts fit. If it does not invest, both earn 0. If it does, the supplier chooses whether to charge a fair price, paying the buyer 6 and the supplier 4, or to exploit the lock-in, paying the buyer -2 and the supplier 8. Find the Nash equilibria and the subgame perfect one.

In the subgame after investment the supplier compares 4 with 8 and exploits. So the buyer, anticipating -2, does not invest, and the unique subgame perfect equilibrium is "do not invest, exploit if invested", paying nothing to either party. The grid also holds a Nash equilibrium in which the buyer does not invest and the supplier plans to charge fairly, which survives only because the plan is never tested, and which subgame perfection discards. What matters is the £10 of joint surplus that vanishes because a promise cannot be made binding, and that is the standard argument for why contracts exist.

Now you. Change the supplier's exploitation payoff from 8 to 3, leaving everything else. What is the subgame perfect outcome now?

Answer

In the subgame after investment the supplier now compares 4 from fair dealing with 3 from exploiting, and deals fairly. The buyer, anticipating 6, invests. The subgame perfect outcome is investment and fair dealing, paying 6 and 4. Nothing about the buyer's problem changed: the investment became safe because the other party's temptation shrank, which is what a reputation, a repeat relationship or a penalty clause is for.

Buying credibility by destroying options

If a threat is not credible because carrying it out would be costly, one repair is to make carrying it out cheap, or to make not carrying it out impossible. The striking thing is that this is worth paying for.

Give the incumbent a prior move. Before the entrant decides, the incumbent may install excess capacity at a cost of 4. The capacity is useless in normal trading but makes a price war cheap to sustain, adding 5 to the payoff from fighting. With the investment made, the payoffs become: staying out pays the incumbent 10-4=6; fighting pays 2+5-4=3; accommodating pays 5-4=1.

Now solve the enlarged game backwards. If the investment has been made, the incumbent facing entry compares 3 with 1 and fights, so the entrant facing an equipped incumbent expects -2 and stays out, leaving the incumbent 6. If the investment has not been made, the earlier analysis applies: the entrant enters, the incumbent accommodates, and the incumbent gets 5. Since 6 beats 5, the incumbent invests, and the market is never entered.

Two features of that calculation are worth isolating.

The investment is a sunk cost by the time entry is decided, and it therefore does not affect the fight-or-accommodate comparison at all: the 4 is subtracted from both branches. What makes the threat credible is the change in the difference between the branches, not the money spent. A commitment that costs a great deal and does not alter any comparison at a future node buys nothing.

And the incumbent is made better off by an action that lowers its payoff in every cell of the table. Before the investment its best outcome was 10 and after it is 6, and it is glad to have made it. That is the central paradox of commitment, and Thomas Schelling's formulation is still the best: the power to constrain an adversary may depend on the power to bind oneself.

Example. What is the most the incumbent would pay for the capacity, and does a higher price make the threat less credible?

Credibility does not depend on the price at all. Whatever k costs, fighting pays 7-k and accommodating pays 5-k after entry, so fighting always wins once the capacity exists. What the price bounds is whether deterrence is worth buying: investing yields 10-k and not investing yields 5, so the incumbent invests when k<5. At k=4 it gains 1. At k=6 it would deter entry and earn 4, which is worse than the 5 it gets by accommodating an entrant, so it does not invest.

Now you. Suppose instead the capacity costs 4 but adds only 2 to the payoff from fighting rather than 5. Does the incumbent invest?

Answer

After entry, fighting would pay 2+2-4=0 and accommodating 5-4=1, so the incumbent would still accommodate. The threat is not made credible, the entrant enters anyway, and the incumbent ends up with 1 instead of 5. The investment is worse than useless: a commitment device that does not reverse the future comparison is money burned. The condition to check is always whether the ranking at the later node has flipped, never how much was spent.

The forms commitment takes

Once the mechanism is clear, the same move can be recognised in very different clothes.

Removing the retreat. Xiang Yu, crossing the river before the battle of Julu in 207 BC, sank the boats and broke the cooking pots. Cortés, at Veracruz in 1519, had his ships scuttled and run aground, a story that later became "burned his boats". Both destroyed their own option to withdraw, which is only rational if the enemy's decision depends on whether withdrawal is possible.

Contracts and penalties. A supply contract with a large penalty for buying elsewhere, a most-favoured-customer clause that forces a firm to refund past buyers if it ever cuts prices, a mortgage that makes walking away expensive. Each turns a preference into an obligation that the other side can verify.

Delegation. Sending an agent whose interests differ from yours: a union negotiator with a mandate, a lawyer paid a contingency fee, a central bank with an inflation target and no instruction to care about employment. A negotiator who genuinely cannot accept less than the mandate is more credible than a principal who could.

Reputation. Doing something costly now so that a later opponent expects the same, which is the subject of the next two lessons.

Brinkmanship. Schelling's insight that a threat too terrible to carry out can be replaced by an action that raises the probability of the terrible outcome without anyone choosing it: massing troops on a border, letting a crisis run, keeping missiles on alert. A commitment that is certain may be incredible while the same commitment made probabilistic is not.

All of these need two properties, and applied analysis should check both. The commitment must be observable, since a commitment the other side cannot verify changes nothing about their reasoning. And it must be irreversible, since one that can be quietly undone is cheap talk with extra steps. Announcements of a pricing policy in a trade journal are commitment because they are both; a private resolution to fight the next entrant is neither.

The chain store paradox

Selten produced the sharpest objection to his own concept in 1978. A chain store operates in twenty towns and faces a potential entrant in each, one after another. The per-town payoffs are those of the entry game.

Solve it backwards. In town twenty there is no future to protect, so the chain accommodates, and the twentieth entrant enters. Knowing that, fighting in town nineteen buys nothing, since the twentieth entrant's decision does not depend on it, so the chain accommodates there too. The argument runs all the way back: the unique subgame perfect equilibrium has entry in all twenty towns and accommodation every time, and the chain earns 20×5=100 rather than the 20×10=200 it would earn by deterring everyone.

Selten's own view was that he would not play this way, and neither would anyone else. Real chains do fight early entrants, and the deterrence appears to work. The paradox is not that the mathematics is wrong; it is that the conclusion depends on twenty layers of confident reasoning about a player who has just been observed doing something the reasoning said was impossible.

Two repairs exist and both matter later. Repetition without a known end changes the arithmetic completely, because there is no last town from which to unravel, and that is the next lesson. Alternatively, keep the finite horizon and add a small doubt: if the entrants think there is even a one per cent chance the chain is the kind that simply enjoys fighting, the chain will fight in the early towns to keep that doubt alive, and deterrence returns. That argument, published in 1982 by David Kreps, Paul Milgrom, John Roberts and Robert Wilson, is one of the most cited results in the subject, and it needs the machinery of private information that arrives near the end of this course.

Example. A chain faces entrants in two towns in sequence, with the entry game payoffs in each. What is the subgame perfect outcome and what does the chain earn in total?

In town two the chain accommodates, so the entrant enters and the chain earns 5 there. In town one, fighting costs the chain 3 (it earns 2 rather than 5) and cannot change what happens in town two, because the town-two subgame has a unique equilibrium regardless of history. So the chain accommodates in town one as well. Both towns are entered, both are accommodated, and the chain earns 10 in total against the 20 it would earn if both entrants stayed out.

Now you. Suppose the entrant in town two is unusual: entering costs it so much that staying out is a dominant strategy there. What does the chain do in town one?

Answer

Town two is settled whatever happens, so it still contributes nothing to the town-one decision, and the chain accommodates in town one exactly as before. The chain earns 5+10=15. The point is that deterrence only pays when the later player's decision actually depends on what is observed, and a later player who is going to stay out anyway is worth nothing to impress.

Where this leaves us

Subgame perfection makes threats accountable, and the price it charges is visible in the chain store: a finite sequence of identical games unravels from the last one, so nothing can ever be sustained by the fear of what comes after.

Real relationships do not have a known last round. A cartel, a marriage, a trade relationship and a border dispute all continue with some probability into a future the participants cannot see the end of. The next lesson puts a number on that future, and finds that the whole conclusion of the prisoner's dilemma reverses once the number is large enough.

Repetition and the shadow of the future

The chain store of the previous lesson unravelled because it had a twentieth town, and almost nothing that matters in trade, politics or family life has a last round anyone can name.

Why a known last round destroys everything

Take the prisoner's dilemma in its numerical form and play it exactly twice. Mutual cooperation pays 4 each, mutual defection pays 2 each, a sole defector takes 6 and leaves the cooperator 1, so in the notation of the dominance lesson T=6, R=4, P=2 and S=1.

Solve it backwards, as the previous two lessons taught. In the second round there is no future to protect. Whatever happened in the first round, the players face a one-shot prisoner's dilemma in which defection strictly dominates, so both defect and each earns 2. That conclusion holds after every possible history, which is the fatal part: the second round pays the same whatever happened in the first. First-round play therefore cannot buy anything, and the first round is itself a one-shot prisoner's dilemma with defection strictly dominant.

The same argument runs for three rounds, for a hundred, for any finite number known to both players. The unique subgame perfect equilibrium of a finitely repeated prisoner's dilemma is defection in every round after every history. Repetition, by itself, achieves precisely nothing.

One caveat is worth stating because it is a genuine exception rather than a technicality. The unravelling needs the stage game to have exactly one Nash equilibrium. If the stage game has two, a good one and a bad one, then the final round can be used as a reward or a punishment, and cooperation becomes sustainable in the earlier rounds of a finite repetition. That is a real construction, and it is also a reminder that the result above is about the prisoner's dilemma in particular, not about finiteness in general.

Putting a number on the future

The escape is to remove the known end. Two quite different things make a payoff next period worth less than the same payoff now, and both are captured by one number.

The first is interest. A pound next year is worth 1/(1+r) today at interest rate r. The second is mortality: the relationship may simply stop. If it continues to the next period with probability 1-p, then next period's payoff arrives only with that probability. Multiply them together and define the discount factor

δ=1-p1+r

so that a payoff x received one period from now counts as δx today. Nothing in the analysis distinguishes the two sources, which is convenient: a relationship that might end and a player who is impatient are the same problem with the same arithmetic.

A constant flow of x per period, starting now and continuing forever, is then worth

x+δx+δ2x+=x1-δ

by the geometric series, which converges because δ<1. The same series has a second reading that is worth carrying around: with continuation probability 1-p and no interest, 1/(1-δ) is the expected number of periods the relationship lasts. A discount factor of 0.9 means an expected ten more rounds; 0.99 means a hundred.

Example. A supplier earns £30,000 a year from a customer. The relationship survives each year with probability 0.9, and the supplier's cost of capital is 5 per cent. What is the relationship worth, and how long is it expected to last?

The discount factor is δ=0.9/1.05=0.8571. The value is 30{,}000/(1-0.8571)=£210{,}000, which is seven years of revenue rather than the twenty a naive undiscounted count of a 10 per cent annual failure rate might suggest. The expected length is 1/(1-0.8571)=7 years, and the two numbers agree because at a zero interest rate the value would be exactly the flow times the expected length.

Now you. The same £30,000 flow, but the relationship survives each year with probability 0.8 and the cost of capital is 4 per cent. Find δ, the value and the expected length.

Answer

δ=0.8/1.04=0.7692, so the value is 30{,}000/(1-0.7692)=£130{,}000 and the expected length is 1/(1-0.7692)=4.33 years. A drop of ten points in the survival probability has cut the relationship's worth by nearly two fifths, which is the sense in which fragile relationships are cheap to betray.

Grim trigger and the patience threshold

Now build a strategy for the indefinitely repeated prisoner's dilemma. The simplest one that works is grim trigger: cooperate in the first round, and cooperate in every later round provided every previous round was mutual cooperation; if anyone ever defects, defect in every round thereafter, forever.

Suppose both players use it. On the path, both cooperate every round, so each earns

R1-δ

Now consider a player who defects in the first round instead. They collect T that round, and from the next round onwards they face a permanent defector, so the best they can do is defect too and collect P forever:

T+δP1-δ

Cooperation is sustained when the first is at least the second. Multiply both sides by 1-δ and the condition collapses to R(1-δ)T+δP, which rearranges to R-Tδ(P-T). Dividing by P-T, a negative number, flips the inequality:

δT-RT-P

Read the fraction. The numerator is what you gain by cheating for one round. The denominator is what you lose, per round, forever afterwards. Patience has to be large enough that the second outweighs the first, and the sucker payoff S never appears, because a player following grim trigger is never caught out by a defection twice.

With the running numbers the threshold is (6-4)/(6-2)=0.5. So a prisoner's dilemma played by players who expect at least one more round as likely as not becomes a game in which cooperation is an equilibrium. Notice what has and has not changed. The stage game still has defection strictly dominant; nobody has become nicer; there is no contract and no enforcement. What changed is that the strategy set now contains plans that condition on history, and one of those plans makes cooperation a best response to itself.

Grim trigger also passes the credibility test of the previous lesson. After a defection, both players defecting forever is the infinite repetition of the stage equilibrium, which is a Nash equilibrium of every subgame it starts. The threat is not merely announced, it is one the punisher is content to carry out.

Example. A stage game has T=7, R=5, P=1 and S=0. What discount factor does grim trigger need, and what expected relationship length does that correspond to at a zero interest rate?

δ(7-5)/(7-1)=2/6=0.3333. At r=0 that means a continuation probability of at least one third, so an expected length of 1/(1-0.3333)=1.5 rounds. Cooperation here is cheap to sustain because the punishment is severe relative to the temptation: cheating gains 2 and costs 4 every round after.

Now you. Another stage game has T=12, R=5, P=2 and S=0. Find the threshold and the expected length it implies.

Answer

δ(12-5)/(12-2)=7/10=0.7, so the expected length must be at least 1/(1-0.7)=3.33 rounds. The temptation is now larger than the whole cooperative payoff, so only a long relationship deters it. Raising T alone always raises the threshold, which is why a one-off windfall available to a partner is the standard way relationships break.

The cartel's own threshold

The quantity-competition lesson left three numbers on the table, computed rather than assumed. In the Cournot market with demand p=100-Q and unit cost £40, the equilibrium gives each firm £400, the shared monopoly gives each £450, and a firm that cheats optimally on the cartel while its rival holds at 15 units earns £506.25. Those are P, R and T.

The threshold follows directly:

δ506.25-450506.25-400=56.25106.25=0.5294

Now do it in symbols, because the answer is more interesting than the arithmetic. Write A=a-c for the gap between the demand intercept and marginal cost. The Cournot profit per firm is A2/(9b), the cartel share is A2/(8b), and the optimal cheat against a rival producing A/(4b) is 3A/(8b) units at a margin of 3A/8, worth 9A2/(64b). Substituting,

δ9/64-8/649/64-1/9=1/6417/576=917=0.5294

Every parameter has cancelled. In any symmetric linear-demand Cournot duopoly with constant marginal cost, the patience needed to hold the full cartel by grim trigger is exactly 9/17, whatever the size of the market, the slope of demand or the cost. That is a genuinely checkable prediction and an uncomfortable one: at a zero termination risk it corresponds to an interest rate of up to 17/9-1=88.9 per cent per period, so any duopoly reviewing prices monthly should collude effortlessly.

Real cartels are not effortless. They form, cheat, break and reform, which means the model is missing something, and the missing piece is usually the assumption that a defection is seen immediately. Suppose cheating takes k periods to detect, so the cheat collects T for k periods before punishment starts. The same algebra gives

δkT-RT-P

and since δ<1, raising k raises the required patience sharply. For the cartel above, a two-period blind spot needs δ9/17=0.7276 and a three-period one needs δ(9/17)1/3=0.809. The first of those is worth checking in full: at δ=0.7276 the cooperative stream is 450/(1-0.7276)=£1{,}652.0, and cheating for two periods before punishment yields 506.25(1+0.7276)+0.72762×400/(1-0.7276)=£1{,}652.0, exactly indifferent.

The moral is that what protects a cartel is not the ferocity of the punishment but the speed of detection, which is why price transparency, most-favoured-customer clauses and published tariffs are treated by competition authorities as collusive devices rather than as conveniences for the customer.

Tit for tat, and the price of forgiveness

Grim trigger works, and nobody would want to use it. One mistaken keystroke, one misread invoice, and a profitable relationship is over forever. The obvious repair is to punish briefly and then forgive, which is tit for tat: cooperate in the first round, then copy whatever the opponent did last round.

Tit for tat needs more patience than grim trigger, and the reason is exactly its forgiveness. Two deviations have to be deterred rather than one.

Against a tit-for-tat opponent, a player who defects forever collects T once and P from then on, which gives the earlier condition δ(T-R)/(T-P). But there is a second and cheaper deviation: alternate. Defect, then cooperate while the opponent retaliates, then defect again as they forgive. That yields T, S, T, S, and so on, worth (T+δS)/(1-δ2). Requiring the cooperative stream to beat it gives R(1+δ)T+δS, that is δ(T-R)/(R-S). Tit for tat therefore sustains cooperation only when

δmax{T-RT-P,T-RR-S}

With T=6, R=4, P=2, S=1 the two thresholds are 0.5 and 2/3, so tit for tat needs δ0.667 where grim trigger needed 0.5. Forgiveness is not free: a punishment that ends is a punishment that can be walked into deliberately.

There is a further honesty about tit for tat that is often skipped. It is not subgame perfect. After a single defection, the two players enter an alternating pattern of defection and cooperation which is not an equilibrium of the subgame it starts, so the strategy relies on carrying out a punishment that the punisher would prefer to abandon. Its reputation comes from tournament performance rather than from the equilibrium arithmetic, and that distinction is the subject of the next lesson.

Example. With T=6, R=4, P=2, S=1 and δ=0.6, does grim trigger sustain cooperation? Does tit for tat?

Cooperating forever is worth 4/(1-0.6)=10. Against grim trigger, defecting yields 6+0.6×2/0.4=6+3=9, which is less than 10, so cooperation holds. Against tit for tat, the alternating deviation yields (6+0.6×1)/(1-0.36)=6.6/0.64=10.3125, which beats 10, so it does not. At δ=0.6 a pair of grim triggers cooperate and a pair of tit-for-tatters are exploited every other round.

Now you. Take T=10, R=6, P=2, S=0 and δ=0.75. Check both strategies.

Answer

Cooperation is worth 6/0.25=24. Permanent defection yields 10+0.75×2/0.25=16, and the alternating deviation yields 10/(1-0.5625)=22.86. Both fall short of 24, so both strategies sustain cooperation here. In threshold form, grim needs (10-6)/(10-2)=0.5 and tit for tat needs the larger of that and (10-6)/(6-0)=0.667, and 0.75 clears both.

What repetition actually needs

The result is strong enough to be misapplied, so it is worth naming the conditions that carry it.

Defections must be observed. Everything above assumes each player sees what the other did. When the signal is noisy, a low price may mean a cheating rival or a weak market, and punishing on the evidence means punishing the innocent. Edward Green and Robert Porter showed in 1984 that the best arrangement under noisy monitoring has price wars occurring on the equilibrium path, triggered by bad demand rather than by anyone's misconduct. Porter's 1983 study of the Joint Executive Committee, the American railroad cartel of the 1880s, found exactly that pattern of periodic collapses in a cartel that was otherwise holding.

There must be no known end. A fixed and commonly known horizon restores the unravelling of the first section. What matters is not that the relationship is literally infinite but that at every point there is a future worth protecting.

The same parties must meet again, and be recognised. Repetition works through identity. Anonymity dissolves it, which is why online markets invest so heavily in making a seller's history follow them and why one-shot tourist trades in every country are the ones with the worst prices.

Nobody has become nicer. Repetition changes the payoffs of the strategies, not the preferences of the players. That also means it is morally neutral: the same argument that sustains a technical standard sustains a price-fixing ring, an omerta, and a cycle of retaliation between neighbours. The theory explains why cooperation is possible, not why it is good.

Where this leaves us

The finite prisoner's dilemma had exactly one equilibrium outcome and it was bad. Adding a shadow of the future has produced a second one, mutual cooperation, above a patience threshold that can be computed from a cartel's own accounts.

It has produced rather more than that, and the excess is a problem. Nothing in the argument was special to mutual cooperation: the same construction, with a suitable punishment behind it, supports alternating exploitation, a ninety-ten split of the gains, or a deliberately wasteful arrangement that neither player likes. The next lesson states exactly how much the repeated game can support, and the answer is close to everything, which turns a triumph into a crisis about what the theory is still predicting.

The folk theorem

The previous lesson showed that patience turns mutual cooperation into an equilibrium of the repeated prisoner's dilemma, and the same construction turns almost every other outcome into one as well.

What repetition can support

Start by drawing the target. In the repeated game a player's payoff is the discounted average of a stream, so it is natural to normalise: multiply the discounted sum by (1-δ) so that a player who earns 4 every period has a repeated-game payoff of 4 rather than 4/(1-δ). On that scale, repeated-game payoffs are directly comparable with stage payoffs.

The set of pairs that could conceivably be achieved is the feasible set: the convex hull of the stage payoff vectors. For the running prisoner's dilemma, with T=6, R=4, P=2 and S=1, the four stage outcomes are (4,4), (6,1), (1,6) and (2,2), and the hull is the quadrilateral they span. Points inside it that are not corners are reached by mixing over time: alternating between (6,1) and (1,6) gives an average of (3.5,3.5), and a rota of three cooperative periods to one exploitative one gives (4.5,3.25). Any point of the hull can be approached this way, and with a publicly observed randomising device, a coin both players see, it can be hit exactly.

Feasibility is not enough to make a point an equilibrium. A player who is being held below what they could get by simply refusing to participate will refuse. So the second ingredient is the floor.

What a player can be held to

The floor is the minmax value: the lowest payoff the other players can force on you when you respond as well as you can to whatever they do.

underline{v}i=mins-imaxsiui(si,s-i)

Read it in that order. The others pick their punishment first; you then best respond; they choose the punishment that leaves your best response worth as little as possible. No equilibrium of the repeated game can ever pay you less than this, because you can always guarantee underline{v}i in every single period by best responding. A payoff vector giving every player strictly more than their minmax is called individually rational.

Two subtleties matter and both are commonly missed. First, minmax is not maximin. Maximin, from the zero-sum lesson, is what you can guarantee when you must move first in the dark; minmax is what others can hold you to when they move first. In a zero-sum game the minimax theorem makes them equal, and in general games they differ.

Second, the minimising side must be allowed to randomise, and this genuinely lowers the floor. Take a game in which player 1 earns 4 from T against L and 0 against R, and 0 from B against L and 3 against R. If player 2 must choose a pure column, player 1 replies with the matching row and earns at least 3, so the pure-strategy floor is 3. If player 2 mixes, playing L with probability q, player 1's best reply is worth max{4q,3(1-q)}, and player 2 minimises that by equalising the two: 4q=3(1-q) gives q=3/7 and a value of 12/7=1.71. Unpredictability is itself a punishment.

In the prisoner's dilemma the minmax is easy. If the opponent cooperates with probability q, your best reply is always to defect, worth 2+4q, which the opponent minimises by never cooperating: underline{v}i=2=P. Mutual defection is exactly the punishment level here, which is why grim trigger was the natural construction.

Example. Find player 1's minmax value in the game where T pays 5 against L and 1 against R, and B pays 0 against L and 4 against R.

Let player 2 play L with probability q. Player 1's best reply is worth max{5q+1(1-q),0q+4(1-q)}=max{1+4q,4-4q}. Player 2 minimises the larger of the two by setting them equal: 1+4q=4-4q gives q=0.375 and a value of 1+1.5=2.5. Against a pure column player 1 could guarantee 4, so mixing costs player 1 a further 1.5.

Now you. In the same game, what would player 1's minmax be if player 2 were restricted to pure strategies, and which column would player 2 use?

Answer

Against L, player 1's best reply is T, worth 5. Against R, it is B, worth 4. Player 2 picks the column with the smaller best reply, so R, holding player 1 to 4. The pure floor of 4 is well above the mixed floor of 2.5, which is why folk theorem statements always minmax in mixed strategies.

The theorem

Put the two ingredients together and the result is the folk theorem, so called because versions of it circulated among game theorists through the 1950s and 1960s with no agreed author.

Take a finite stage game with minmax values underline{v}i. For any feasible payoff vector v with vi>underline{v}i for every player, there is a discount factor δ<1 such that for every δ>δ, v is the payoff of a subgame perfect equilibrium of the infinitely repeated game.

The construction is the one already seen. Put the players on a path that yields v, and specify that any deviation triggers a punishment phase in which the deviator is held to their minmax. Deviating gains at most a bounded amount in the period it happens and costs a stream, so once δ is high enough the stream wins.

Two versions are worth separating. James Friedman proved in 1971 that any feasible payoff strictly better for everyone than a Nash equilibrium of the stage game can be sustained by reversion to that equilibrium. This is the easy and completely credible version, since the punishment is a Nash equilibrium and nobody has to be persuaded to carry it out. Drew Fudenberg and Eric Maskin proved in 1986 that the stronger statement, with the punishment level pushed all the way down to the minmax, also gives subgame perfect equilibria, provided the feasible set has full dimension. The extra work in their proof is entirely about making the punishers willing to punish, since minmaxing someone can be costly for the punisher, and they solve it by rewarding the punishers afterwards.

Building one, and finding it is not cooperation

Take the running prisoner's dilemma and aim at (3.5,3.5), achieved by taking turns at exploitation: in even periods player 1 defects while player 2 cooperates, in odd periods the reverse. Any departure from this rota triggers mutual defection forever.

The binding constraint belongs to whoever is due to be the victim this period, because they are the one asked to accept 1 now in exchange for 6 next time. Following the rota is worth

S+δT1-δ2=1+6δ1-δ2

Deviating means defecting now and facing permanent mutual defection, worth 2/(1-δ). Setting the first at least equal to the second, and cancelling a factor of 1-δ, gives 1+6δ2(1+δ), so δ0.25.

That is a strikingly low bar, and the outcome it supports pays each player 3.5 per period, less than the 4 they would get from plain cooperation and more than the 2 of the stage equilibrium. So for any δ above 0.25 the repeated game has an equilibrium in which the players systematically take turns at exploiting each other, wasting half a point each, and neither can improve by deviating. Nothing marks it as worse than mutual cooperation from inside the model, because equilibrium is a statement about deviations, not about efficiency.

Example. Suppose mutual defection paid 3 instead of 2, leaving everything else unchanged. What happens to the threshold for the alternating scheme?

The victim's condition becomes 1+6δ3(1+δ), so 3δ2 and δ0.667. A punishment worth 3 rather than 2 is a weaker punishment, so more patience is needed to make the same arrangement stick. Both numbers are still below the minmax check, since 3.5 exceeds the new floor of 3, so the arrangement remains individually rational, but only just.

Now you. In the original game, does the rota still work if the punishment is not permanent but lasts a single period of mutual defection, after which the rota resumes where it left off?

Answer

The victim's gain from deviating is 2-1=1 this period, and the cost is one period of mutual defection instead of one period of exploiting, worth δ(6-2)=4δ, plus nothing after that since the rota resumes. So the condition is 4δ1, the same δ0.25 as before. Shortening the punishment did no damage here because the deviation gain is small; against a larger temptation the shorter punishment would fail first, which is the tit for tat calculation of the previous lesson in another guise.

A theory that predicts everything

The folk theorem is usually presented as the triumphant explanation of cooperation, and it is more accurate to call it the moment the theory stops predicting.

Count what has been established. The repeated prisoner's dilemma has equilibria at mutual cooperation, at mutual defection, at alternating exploitation, at a ninety-ten division of the gains, at deliberately wasteful arrangements that nobody likes, and at every other individually rational point of a two-dimensional region. There are uncountably many equilibrium payoff vectors, and any observed pattern of long-run behaviour between two parties can be presented as one of them.

That is not a prediction. Contrast it with the zero-sum lesson, where the theory named a single number, said kickers should go left 38.5 per cent of the time, and could have been wrong. Here nothing could be observed that would embarrass the model, and a claim that cannot fail is doing no work. The multiplicity problem raised in the coordination lesson has not merely persisted, it has gone from three equilibria to a continuum.

Three responses exist and none of them is fully satisfying. One is to add a selection criterion from outside, as with focal points and risk dominance. One is to charge for complexity: Dilip Abreu and Ariel Rubinstein showed in 1988 that if strategies must be implemented by finite automata and simpler machines are cheaper, the set of equilibria shrinks sharply. The third is to stop deriving and start measuring, which is where the subject actually went.

Axelrod's tournaments

Robert Axelrod's move in 1980 was to treat the question as empirical rather than deductive. He invited people to submit computer programs to play the repeated prisoner's dilemma, ran a round robin, and asked which did well against the field that had actually turned up.

The first tournament had fourteen entries plus a random player, each pairing lasting 200 rounds, with the scoring T=5, R=3, P=1, S=0. The second, publicised on the results of the first, drew sixty-two entries from six countries, and to remove the end-game problem the length was made probabilistic, with a continuation probability of 0.99654 giving a median match of 200 moves.

Tit for tat won both. It was submitted by the psychologist Anatol Rapoport and was the shortest program entered, four lines of BASIC. Axelrod's summary of why is worth quoting in his own categories. It was nice, never the first to defect. It was provocable, retaliating immediately. It was forgiving, returning to cooperation as soon as the opponent did. And it was clear, simple enough for an opponent to work out what it was doing and therefore to see that cooperation paid.

He then ran an ecological version of the second tournament, repeating it over generations with each strategy's population share growing in proportion to its score. Tit for tat's share rose steadily. Strategies that exploited the naive did well in the first generations and then starved as their prey died out, which is a genuinely interesting mechanism and not one that any equilibrium argument had suggested.

Example. In the tournament scoring, what do tit for tat and unconditional defection score against each other over 200 rounds?

Round one, tit for tat cooperates and is defected on: 0 against 5. From then on tit for tat defects too, so both take 1 for the remaining 199 rounds. Tit for tat scores 199 and the defector scores 204. This is the fact most often left out of the popular retelling: tit for tat never beat any opponent head to head in either tournament. It won by never losing badly, while its rivals ruined each other.

Now you. Two tit for tat players meet over 200 rounds, but in round 100 one player's intended cooperation comes out as a defection. What does each score, against the 600 they would have scored without the error?

Answer

Rounds 1 to 99 are mutual cooperation, 297 each. In round 100 the slipped move scores 5 to the accidental defector and 0 to the other. From round 101 the two are locked in alternating retaliation, each scoring 5 in half of the remaining 100 rounds and 0 in the other half, so 250 each. Totals are 552 and 547 against 600. One slipped move in two hundred costs each player about a twelfth of their score, and the echo never dies out.

What the tournaments do and do not show

A tournament result is not a theorem, and the qualifications are substantial.

The winner depends on the field. Tit for tat won against the strategies people submitted; a field stuffed with unconditional defectors would have been won by unconditional defection, since tit for tat's advantage comes entirely from meeting other cooperative strategies. Axelrod was clear about this, and later work made it sharper: Robert Boyd and Jeffrey Lorberbaum showed in 1987 that no strategy at all is evolutionarily stable in the repeated prisoner's dilemma, so there is no final winner to be found.

Noise is worse for tit for tat than the tournaments suggested, as the exercise above shows. Two repairs do better under noise. Generous tit for tat forgives a proportion of defections at random, breaking the echo. Win-stay lose-shift, called Pavlov, repeats its last move if the outcome was good and switches if it was bad, which lets it recover from errors and also lets it exploit an unconditional cooperator, something tit for tat never does. Martin Nowak and Karl Sigmund's 1993 simulations found Pavlov displacing tit for tat once errors are allowed.

And there are strategies nobody had thought of. William Press and Freeman Dyson showed in 2012 that a class of memory-one strategies, which they called zero-determinant, can unilaterally fix a linear relationship between the two players' scores, allowing an extortionate player to guarantee themselves a fixed multiple of the opponent's surplus. Against an opponent who adapts, extortion works. Against another extortionist, both do badly. That result appeared thirty-two years after the tournaments and shows how far from closed the question was.

Where this leaves us

Repetition explains cooperation and rather too much else besides. The honest summary is that the shadow of the future makes cooperation possible without making it necessary, that which arrangement a particular pair settles into is not determined by the payoffs, and that the empirical work points to simple, retaliatory, forgiving rules rather than to anything the equilibrium arithmetic singled out.

That leaves an obvious gap. If a long relationship can support any division of the gains, what actually decides the division? The question is sharpest when the division is the whole point: two parties with a surplus of £100 to share and no way to create it except by agreeing. Equilibrium accepts every split from 0 to 100, so it says nothing at all. The next lesson takes the opposite approach and asks what properties a fair and sensible split ought to have, then shows that four modest-sounding properties pin down exactly one answer.

Splitting the surplus

When two parties have a hundred pounds to divide and no way to create it except by agreeing, equilibrium analysis accepts every division from nothing to everything, which is not a prediction but a shrug.

The problem equilibrium cannot solve

Write the division as a game. Both parties simultaneously name a demand, x for the first and y for the second. If x+y100 each receives what they asked for; otherwise negotiations fail and both receive nothing.

Now hunt for Nash equilibria. Suppose the second party demands 40. The first party's best reply is 60: demanding more means the total exceeds 100 and both get nothing, demanding less means leaving money on the table. So (60,40) is an equilibrium. The identical argument works for (50,50), for (99,1), for (1,99) and for every other exactly exhausting pair. There is even an equilibrium at (100,100), where both demand everything, both get nothing, and neither can improve alone, since a unilateral reduction still leaves the total above 100.

So the game has a continuum of equilibria, one for every split, plus a disastrous one. Refinement does not help: there are no unreached nodes, so subgame perfection has nothing to remove, and the folk theorem of the previous lesson says the same problem returns in worse form if the parties bargain repeatedly. The equilibrium concept is not merely failing to choose here. It is structurally incapable of choosing, because every split is a mutual best response and that is all the concept looks at.

John Nash, then twenty-one, published a different move in Econometrica in 1950. Stop deriving the split from individual optimisation. Instead state the properties a sensible answer ought to have, and see how many answers survive.

Writing the problem down

A bargaining problem is a pair (S,d). The feasible set S is the set of utility pairs (u1,u2) the parties can achieve by agreement, taken to be convex, closed and bounded above. The disagreement point d=(d1,d2) is the utility pair they receive if they fail to agree.

Both objects need care. The feasible set is in utilities, not in pounds, which matters as soon as the parties are not equally happy about risk. And the disagreement point is not zero by convention; it is whatever each party actually gets from walking away. A union with a strike fund and a firm with six months of stock have a different disagreement point from a union with no savings and a firm with none, and it will turn out that this is the single most important number in the whole problem.

Convexity is worth a sentence, since it is doing real work. If two agreements are available, the parties can agree to flip a coin between them, and if payoffs are von Neumann-Morgenstern utilities, as the first lesson insisted, the value of that coin flip is the average of the two. So the feasible set automatically contains every point on the line between any two of its points.

Four properties

Nash asked for a rule f that takes any bargaining problem (S,d) and returns a single point of S, subject to four requirements.

Efficiency. The chosen point is on the Pareto frontier: there is no other feasible point that is at least as good for both and strictly better for one. Leaving money unclaimed is not an answer.

Symmetry. If the problem is symmetric, meaning d1=d2 and S is unchanged when the two coordinates are swapped, then the solution gives both parties the same utility. This says the rule uses nothing about the parties except what is in S and d, so it cannot favour the taller or the older or the one whose name comes first.

Invariance to affine transformations. If one party's utility scale is rescaled, u1αu1+β with α>0, the solution is the same agreement, described in the new units. This is the invariance from the first lesson, and it is what prevents a party from improving their share by writing their utilities in different units.

Independence of irrelevant alternatives. If the solution to a problem lies inside a smaller feasible set with the same disagreement point, then it is the solution to the smaller problem too. Removing options that were not going to be chosen changes nothing.

None of these sounds like it says much. Together they say everything.

One answer survives

Theorem (Nash, 1950). There is exactly one rule satisfying all four, and it selects the point of S with uidi that maximises the Nash product

(u1-d1)(u2-d2)

The proof of uniqueness is a two-line trick worth knowing. Given any problem, use the affine invariance axiom to rescale both utilities so that the Nash product maximiser sits at (1,1). One can then show the whole feasible set lies below the line u1+u2=2, since a point above it would allow a higher product. Enlarge the problem to the symmetric triangle under that line: by symmetry and efficiency, the rule must choose (1,1) there. By independence of irrelevant alternatives, since (1,1) is in the original smaller set, it must be chosen in the original problem too. Undo the rescaling and the theorem is proved.

Apply it to money first, where it is almost too simple. With linear utilities, a pot of M and a disagreement point (d1,d2), the frontier is u1+u2=M and the product (u1-d1)(M-u1-d2) is a downward parabola in u1 with roots at d1 and M-d2. Its peak is halfway between them:

u1=d1+M-d1-d22

So the Nash solution takes each party's outside option off the top and splits the remaining surplus down the middle. That is not an assumption; it is what four axioms about fairness and consistency force.

Example. Two firms can jointly earn £100,000 from a project. Without the deal, firm 1 has an alternative worth £30,000 and firm 2 has nothing. Both are risk neutral. What does the Nash solution give?

Maximise (x-30)(100-x) in thousands. The roots are 30 and 100, so the peak is at x=65. Firm 1 takes £65,000, firm 2 takes £35,000, and each has gained £35,000 over its outside option. The £30,000 alternative was worth exactly £30,000 in the bargain, no more and no less, which is the clearest reason for a negotiator to invest in an alternative before negotiating rather than in rhetoric during.

Now you. Same £100,000, but firm 1's outside option is £20,000 and firm 2's is £5,000. Find the split.

Answer

Maximise (x-20)(95-x), whose roots are 20 and 95, so x=57.5. Firm 1 takes £57,500 and firm 2 takes £42,500, and each gains £37,500 over its own fallback. The surplus being divided is 100-20-5=£75{,}000, split evenly, with the fallbacks returned on top.

When the frontier is not a straight line

Not every bargain trades pound for pound. If one party's pound costs the other two, because of tax, transport or a difference in what the asset is worth to each, the frontier tilts and the Nash solution moves with it.

Example. A frontier runs x+2y=120, so every pound given to party 2 costs party 1 two pounds. The disagreement point is (0,0) and both are risk neutral. Where does the Nash solution sit?

Maximise xy subject to the constraint. Substituting y=(120-x)/2 gives the product x(120-x)/2, a parabola with roots 0 and 120, peaking at x=60. So x=60 and y=30, with a Nash product of 1800. Party 1 takes twice as much in nominal terms, and in utility terms the two are equally distant from disagreement given the exchange rate between them. The general pattern for a linear frontier through the origin is that each party takes half of what they could have taken alone.

Now you. The frontier is x+3y=90 with disagreement at (0,0). Find the Nash solution.

Answer

Maximise x(90-x)/3, roots 0 and 90, peaking at x=45, so y=15 and the Nash product is 675. Party 1 could have taken 90 alone and takes 45; party 2 could have taken 30 alone and takes 15. Both take half their maximum, which is what a linear frontier through the disagreement point always gives.

What risk aversion costs

Now the point where working in utilities rather than money stops being pedantry.

Two people split £100. The first has utility u1(x)=x, so they are risk averse in the sense of the first lesson. The second is risk neutral, u2(y)=y. Disagreement pays both nothing. Maximise the Nash product x(100-x).

Taking logs turns the product into 12lnx+ln(100-x), and setting the derivative to zero gives 12x=1/(100-x), so 100-x=2x and x=33.33. The risk-averse party receives £33.33 and the risk-neutral one £66.67.

The general version is worth having. If u1(x)=xα with 0<α1 against a risk-neutral opponent, the same calculation gives

x=100α1+α

At α=1 the parties are equally risk neutral and split evenly. At α=0.5 the share is £33.33, at α=0.25 it is £20. The more sharply diminishing your marginal utility, the less you get.

The mechanism is not that risk aversion makes you a worse negotiator by temperament. It is that a bargaining problem is implicitly a problem about the risk of disagreement, and a party who suffers more from a bad outcome is willing to concede more to avoid it. That is a real and testable prediction about who does badly in negotiations, and it is also the sharpest warning against putting money in the payoff cells: had both parties' payoffs been written in pounds, the split would have come out at fifty each and the prediction would have been wrong.

Example. The same £100 with u1(x)=x, but now the risk-averse party has an outside option worth £16 while the other has nothing. What is the split?

The disagreement point in utilities is (16,0)=(4,0), so maximise (x-4)(100-x). The derivative of the log gives 12x(x-4)=1100-x, so 100-x=2x-8x. Writing s=x gives 3s2-8s-100=0, with positive root s=(8+64+1200)/6=(8+35.55)/6=7.259, so x=52.7. The outside option was worth £16 and it bought £19.4 of extra share, because it also removed part of the risk that made the party concede.

Now you. Back to no outside options, but with u1(x)=x0.25 against a risk-neutral partner. What does the risk-averse party get?

Answer

x=100(0.25)/1.25=£20. Compared with the £33.33 at α=0.5 and the £50 at α=1, the pattern is monotone: every increase in risk aversion transfers share to the other side. Nothing about the two parties differs except the curvature of one utility function.

The axiom that is doing the damage

Three of the four axioms are close to unarguable. The fourth, independence of irrelevant alternatives, is where the objections live, and the objection is easy to state with numbers.

Two parties split £100 with linear utilities and no outside options, so the Nash solution is £50 each. Now suppose a regulator caps party 2's share at £60. The cap removes only agreements that were not going to be chosen, so by independence of irrelevant alternatives the solution stays at £50 each. And yet something real has changed: party 1 can still imagine receiving the whole £100 while party 2's best case has fallen to £60, so their positions are no longer symmetric in any intuitive sense.

Ehud Kalai and Meir Smorodinsky proposed in 1975 dropping independence and requiring monotonicity instead: if the frontier expands so that one party's best possible outcome rises with the other's held fixed, that party should not do worse. Their solution equalises the ratio of each party's gain to their maximum possible gain. In the capped example the ideal point is (100,60), so the solution sits where x/100=y/60 on the frontier x+y=100. Writing both as t times their ideal gives 160t=100, so t=0.625 and the split is £62.50 to party 1 and £37.50 to party 2. The cap on party 2's upside has cost them £12.50 in a bargain they were never going to push that far.

Which is right is not settled and probably has no answer in the abstract, because they are answers to different questions. Nash's rule is the one that is consistent under removing options; Kalai and Smorodinsky's is the one that responds to how much each party could in principle have got. Experimental subjects, offered problems where the two disagree, do not systematically match either.

One further honesty is owed. Nash's theorem selects a point but describes no process. Nobody in it makes an offer, refuses one, or walks out. It is a statement about what a reasonable arbitrator would rule, and it is silent about two people in a room with conflicting interests and time to waste. Nash was aware of the gap and proposed closing it by building non-cooperative models whose equilibria reproduce the axiomatic answer, a research programme now called the Nash program.

Where this leaves us

Four modest requirements pick out one split from a continuum that equilibrium could not narrow at all. The split hands each party their outside option and divides what is left in half, shifting against whoever is more risk averse and whoever has less to fall back on.

That is a satisfying answer to the wrong question. Real bargaining is a sequence: someone opens, someone counters, time passes, and every day of delay costs both sides money. The next lesson models the haggling itself, with impatient players making alternating offers, and finds that the procedure has a unique subgame perfect outcome which converges on the split derived here as the delay between offers shrinks. It also finds that people in laboratories do not play it.

Bargaining by taking turns

The previous lesson picked a split by asking what properties a good answer should have, and it described nobody making an offer, refusing one, or storming out.

One offer, take it or leave it

Start with the shortest possible negotiation. Party A proposes a division of £100 and party B either accepts, in which case it happens, or rejects, in which case both get nothing. This is the ultimatum game, and backward induction settles it in a line.

At the final node B compares the offer with zero. Any positive amount beats nothing, so B accepts anything above zero. Anticipating that, A offers the smallest positive amount and keeps the rest. If money is divisible into pennies the unique subgame perfect equilibrium is A keeping £99.99, and if it is perfectly divisible the equilibrium is A keeping everything with B indifferent and accepting.

The result is not an artefact of the tiny horizon; it is the whole logic of commitment from the earlier lesson, arrived at for free. A has the power to make a take-it-or-leave-it offer, which is a commitment device the rules handed over, and the entire surplus follows the commitment. What makes the prediction interesting is that it is so easy to test, and it fails badly, which the second half of this lesson takes up.

Letting the other side answer back

Give B a right of reply. If B rejects A's offer, a period passes and B makes a counter-offer, which A may accept or reject; if A rejects, both get nothing. Delay is costly: a pound agreed one period later is worth δ pounds now, with δ the discount factor of the repetition lesson.

Solve backwards. In the second period B is making a take-it-or-leave-it offer, so B takes the whole pie, worth δ to B in first-period terms. So in the first period B will accept anything worth at least δ, and A offers exactly that, keeping 1-δ. With δ=0.9, A keeps just 10 per cent. Having the last word is worth almost everything when players are patient.

Now add a third period in which A proposes again. In period three A takes everything, so in period two B must leave A a share worth δ, keeping 1-δ=0.1. In period one A must leave B a share worth δ(1-δ)=0.09, keeping 0.91.

The shares as the horizon grows, still at δ=0.9, run 1, 0.1, 0.91, 0.181, 0.8371, 0.2466, 0.7781, and so on. They oscillate, because whoever holds the last offer takes everything and each extra period flips who that is, and they converge, because the swings shrink by a factor of δ each time. The limit is 0.5263.

Example. With δ=0.8 and three periods of alternating offers, what does the first proposer keep?

In period three the first proposer takes everything. So in period two the second party must leave them δ=0.8, keeping 0.2. In period one the first proposer must leave the other δ(0.2)=0.16, so keeps 0.84.

Now you. Same δ=0.8, but with four periods. What does the first proposer keep, and why has the answer moved so far?

Answer

Working back: the fourth-period proposer is the second party, who takes 1. The third-period proposer keeps 1-0.8=0.2. The second-period proposer keeps 1-0.8(0.2)=0.84. The first-period proposer keeps 1-0.8(0.84)=0.328. The answer swung from 0.84 to 0.328 because an even number of periods hands the last offer to the other side, and the last offer is the commitment that everything else discounts back from.

Never having the last word

Ariel Rubinstein's 1982 paper removed the last word entirely. Offers alternate indefinitely: A proposes, B accepts or counters, A accepts or counters, forever, with each period of delay shrinking the pie by the factor δ. There is now no final node, so backward induction has nowhere to start.

Use stationarity instead. Every subgame that begins with a given player proposing looks exactly like every other, so if a player's equilibrium share when proposing is x, it is x in every such subgame. Then the responder's choice is between accepting what they are offered now and rejecting to become the proposer one period later, which is worth δx to them. The proposer offers exactly δx, no more, and keeps the rest:

x=1-δxx=11+δ

At δ=0.9 the proposer keeps 0.5263 and the responder 0.4737, matching the limit of the finite sequence. At δ=0.5 the proposer keeps 2/3. As δ rises to 1 the split approaches even, and the proposer's advantage, (1-δ)/(1+δ) of the pie, is exactly the cost of one period's delay.

Rubinstein's real achievement was uniqueness rather than the formula. Stationarity was assumed above; his proof does not assume it. Let M be the highest and m the lowest share a proposer receives in any subgame perfect equilibrium. A responder can always secure δm by rejecting, so no proposer keeps more than 1-δm, giving M1-δm. A responder will never turn down more than δM, so every proposer can secure at least 1-δM, giving m1-δM. The two inequalities force M=m=1/(1+δ). The equilibrium is unique among all subgame perfect equilibria, which is a much rarer outcome than the previous lessons might suggest.

Two features of the solution deserve attention because they are testable. Agreement is immediate: the first offer is accepted, and no delay is ever observed on the equilibrium path. And the shrinking of the pie is never actually paid; it works entirely through the threat of paying it, in the same way that the commitment lesson's capacity investment worked through a comparison that was never reached.

Example. Two firms bargain over £100,000 of joint surplus, alternating offers monthly, and each values money a month later at 0.9 of its present worth. What is the equilibrium division?

The proposer keeps 1/(1+0.9)=0.5263, so £52,632, and the responder takes £47,368. The proposer's advantage is £5,264, which is 10 per cent of the £52,632 the responder would be waiting a month for, and that is the entire content of "first-mover advantage" here.

Now you. The same firms bargain annually rather than monthly, so that a year's delay leaves them valuing money at 0.5 of its present worth. What is the division now?

Answer

The proposer keeps 1/1.5=2/3, so £66,667 against £33,333. Slowing the rounds down without changing anything else has transferred £14,000 to whoever speaks first, which is why the frequency with which offers may be made is itself worth negotiating over before the negotiation starts.

Impatience decides

Nothing forces the two parties to be equally patient, and once they are not, the formula says something with real content.

Let party 1 discount at δ1 and party 2 at δ2. Write x for party 1's share when party 1 proposes and y for party 2's share when party 2 proposes. Party 1, proposing, must leave party 2 the discounted value of party 2's own proposer share, so x=1-δ2y. Symmetrically y=1-δ1x. Substituting the second into the first gives x=1-δ2+δ1δ2x, so

x=1-δ21-δ1δ2

With δ1=0.9 and δ2=0.8, party 1 receives 0.2/0.28=0.7143. Patience is worth 71 per cent of the pie against an opponent who loses a fifth of the value with every round of delay. Swap the two and party 1's share falls to 0.3571.

The comparative static is the useful part. Your share rises with your own patience and falls with your opponent's, and the mechanism is entirely about who can afford to wait: an impatient party is one for whom every rejection is expensive, and the other side prices that. This is what a strike fund buys, what a firm's inventory buys, and what a mortgage payment due next week takes away.

Example. A union discounts at δ1=0.95 per week and the firm at δ2=0.7 per week, because the firm's plant is idle and the union has a strike fund. The union proposes first over a £2 million surplus. What does the union get?

x=(1-0.7)/(1-0.665)=0.3/0.335=0.8955, so £1.79 million. Almost the entire surplus goes to the side that can wait, and the firm's weekly loss of 30 per cent of the value is what pays for it.

Now you. Suppose the firm builds up stock so that its weekly discount factor rises from 0.7 to 0.9, and the union's stays at 0.95. What is the union's share now?

Answer

x=(1-0.9)/(1-0.855)=0.1/0.145=0.6897, so £1.38 million. The stockpile is worth £410,000 to the firm, and it is worth that without a single day of strike actually happening. Preparing to endure a dispute changes the terms of the settlement that avoids it, which is the commitment lesson again in a different suit.

Meeting the axioms in the limit

Two apparently unrelated answers to the same question now sit side by side, and they turn out to be the same answer.

Let the interval between offers be Δ and let each party have a continuous discount rate ri, so δi=e-riΔ. As Δ shrinks, 1-δiriΔ, and substituting into the formula gives

xr2r1+r2

The share depends only on the ratio of impatience, and the first-mover advantage vanishes, because the advantage was worth one period's delay and the period has become negligible. Numerically, with r1=5 per cent and r2=15 per cent, the discrete formula gives 0.768 at Δ=1, 0.752 at Δ=0.1 and 0.7502 at Δ=0.01, converging on the predicted 0.75.

That limit is exactly the asymmetric Nash bargaining solution of the previous lesson, with bargaining weights r2/(r1+r2) and r1/(r1+r2). Ken Binmore, Ariel Rubinstein and Asher Wolinsky set this out in 1986, and the moral is the one Nash had hoped for: an axiomatic answer earns its keep when a procedural model of the haggling delivers it. When it is the disagreement point rather than impatience that drives the outcome, a different procedural model converges on the plain Nash solution with the outside options in the disagreement point, so the axioms are not one thing but a family whose members depend on what actually threatens the negotiation.

What people actually do

The ultimatum game is the easiest prediction in this course to test, and it is comprehensively wrong.

Werner Güth, Rolf Schmittberger and Bernd Schwarze ran it in 1982. The theory says offer the smallest positive amount and have it accepted. What happens, across hundreds of replications since, is that the modal offer is 40 to 50 per cent of the pie, mean offers cluster around 40 per cent, and offers below about 20 per cent are rejected roughly half the time. Rejection means choosing nothing over something, and it persists when the stakes are raised: studies paying two or three months' income in Indonesia and in India still find low offers refused.

The cross-cultural evidence sharpens rather than dissolves the puzzle. Joseph Henrich and colleagues ran the game in fifteen small-scale societies in 2001 and found the variation between societies far larger than anything within industrialised samples. The Machiguenga of Peru, who live in largely independent households, offered about 26 per cent on average and almost never rejected. The Lamalera whale hunters of Indonesia, whose livelihood requires large cooperative crews, offered about 58 per cent. Among the Au and Gnau of Papua New Guinea, where accepting a gift creates an obligation, offers above half were rejected as often as offers below it. The prediction fails in a different direction in each place, which is evidence that the payoffs in the laboratory are not the payoffs in the participants' heads.

Alternating offers fares no better. Jack Ochs and Alvin Roth's 1989 experiments found first offers far from the subgame perfect prediction, and, more damagingly, found that most rejections were followed by a disadvantageous counter-offer: the rejecting party countered with a demand worth less to themselves, after discounting, than what they had just turned down. No adjustment to the discount factor rescues that, since it is not a mistake about arithmetic but a refusal to be treated a certain way.

What the model still buys

The natural conclusion, that bargaining theory is useless, is too quick. Three things survive.

The comparative statics survive. Whatever the level of offers, the direction of the effects is right: the more patient side does better, better outside options improve terms, and the ability to make the final offer is worth something. These predictions hold in the experiments even where the point predictions do not.

The mechanism survives. Impatience is a real cost of disagreement, and the model correctly identifies that the cost of failing to agree, rather than any notion of deservingness, is what moves the split. Anyone preparing for a negotiation who spends their effort on their alternative rather than on their argument is applying it correctly.

And the failure is informative. The model predicts immediate agreement, and real disputes involve strikes, lockouts, delayed settlements and litigation that both sides know will be settled eventually. The best explanation for delay is the assumption this course has not yet dropped: that each side knows what the other's payoffs are. If a party's patience or reservation value is private, refusing an offer becomes a way of signalling that you are the tough type, and delay stops being a mistake and becomes information transmitted at a price.

Where this leaves us

Two routes to the same split have now been travelled, one axiomatic and one procedural, and they converge as the delay between offers goes to zero. Both assume that each party knows the other's preferences exactly.

That assumption has been in force since the first lesson, and it is false in most of the situations the subject is applied to. A bidder does not know what a painting is worth to the person beside them, an insurer does not know how careful a customer is, and a negotiator does not know how badly the other side needs a deal. The next lesson gives each player a private type and a belief about everyone else's, which is the last major piece of machinery in the course, and applies it to the setting where the theory has done most practical good: auctions.

Private information and auctions

Every model so far has assumed each player knows the other's payoffs exactly, and a bidder at an auction does not know what the object is worth to the person beside them.

Giving each player a secret

The difficulty is easy to state and looks fatal. If you do not know my payoffs, you cannot compute my best response, so you cannot compute yours. Worse, you have beliefs about my payoffs, and I have beliefs about your beliefs, and there is no obvious place for that regress to stop.

John Harsanyi's solution, in a three-part paper published in Management Science across 1967 and 1968, was to convert the unknown into a move by nature. Each player has a type θi, a private variable holding everything they know and others do not: their valuation, their cost, their patience. Nature draws the whole profile of types from a probability distribution that is common knowledge, and then tells each player their own type and nothing else. The regress collapses, because everyone's beliefs about everyone's beliefs are derived by conditioning the same shared prior, rather than being separate objects to be specified.

The move is not free. It replaces "I do not know your payoffs" with "I do not know your payoffs but I know exactly the distribution they were drawn from", which is a substantial assumption and is precisely what the common prior gives up. In an auction for oil rights it is defensible; in a first meeting between two firms from different industries it is a fiction that should be flagged.

A strategy in such a game is no longer a single action. It is a function from types to actions: a rule saying what you do for each valuation you might turn out to have. A Bayes-Nash equilibrium is a profile of such rules in which each type of each player is maximising expected payoff, the expectation taken over the other players' types using the prior and over their actions using their rules. This is the same equilibrium idea as ever, applied to a game in which the strategies are functions.

The auction as a laboratory

An auction is the setting where all of this pays off, for three reasons: the type is one number, the rules are written down, and the outcomes are observed in public with money attached.

Fix the standard independent private values model. There is one indivisible object and n bidders. Bidder i's value vi is drawn independently from the uniform distribution on [0,1], is known to bidder i alone, and is unaffected by what anyone else's value turns out to be. A bidder who wins at a price p gets vi-p and a bidder who loses gets nothing.

Four formats are standard. In an English auction the price rises until one bidder remains. In a Dutch auction the price falls until someone claims the object. In a first-price sealed bid the highest bid wins and pays what it bid. In a second-price sealed bid, proposed by William Vickrey in 1961, the highest bid wins and pays the second-highest bid.

Truth-telling in a second-price auction

The second-price rule looks like a mistake, since the seller deliberately collects less than the winner offered. It buys something in return.

In a second-price auction, bidding your true value weakly dominates every other bid. The argument needs no probabilities at all, which is what makes it so strong. Let h be the highest bid among your rivals, whatever it is.

Consider bidding above your value v. This changes nothing unless h lands between v and your bid. In that case you win and pay h>v, taking a loss, where truthful bidding would have lost the auction and paid nothing. Now consider bidding below v. This changes nothing unless h lands between your bid and v. In that case you lose the auction, where truthful bidding would have won it at a price below your value. Every deviation is either irrelevant or harmful, which is weak dominance exactly as defined in the dominance lesson.

The consequence is that a second-price auction requires no strategic thought whatsoever. A bidder does not need to estimate how many rivals there are, how aggressive they are, or how the values are distributed. This is the property that makes the format attractive to designers, and it is the same argument that runs an English auction, where staying in until the price reaches your value is optimal for the same reason.

Example. Your value for a painting is £70 in a second-price sealed-bid auction. Compare bidding £70, £85 and £55 against a highest rival bid of first £80 and then £60.

Against £80: bidding 70 loses, worth 0; bidding 85 wins at a price of 80, worth 70-80=-£10; bidding 55 loses, worth 0. Against £60: bidding 70 wins at 60, worth £10; bidding 85 also wins at 60, worth £10; bidding 55 loses, worth 0. Truthful bidding is never beaten and it strictly beats overbidding in the first case and underbidding in the second.

Now you. Your value is £40 and you bid £50. The highest rival bid turns out to be £45. What happens, and what would truthful bidding have given?

Answer

You win and pay £45 for something worth £40 to you, a loss of £5. Truthful bidding would have lost the auction and paid nothing, which is better. Overbidding only ever changes the outcome in the range where winning is a mistake, and that is the whole content of the dominance argument.

Shading in a first-price auction

Now the format where thinking is required. In a first-price sealed-bid auction, bidding your value guarantees a payoff of zero whether you win or lose, so every bidder shades downwards. How far depends on how many rivals there are, which means the equilibrium has to be solved for.

Look for a symmetric equilibrium in which every bidder uses the same increasing rule, and guess that it is linear: b(v)=kv for some constant k to be found. Take one bidder with value v considering a bid b. They win when every one of the other n-1 bidders has kvj<b, that is vj<b/k. Since values are uniform on [0,1], each such event has probability b/k, and independence multiplies them:

expected payoff=(v-b)(bk)n-1

The constant kn-1 does not affect where the maximum is, so maximise (v-b)bn-1. Differentiating and setting to zero gives -bn-1+(v-b)(n-1)bn-2=0, and dividing through by bn-2 leaves -b+(n-1)(v-b)=0, so nb=(n-1)v and

b(v)=n-1nv

The guess is confirmed, with k=(n-1)/n. Read the formula. With two bidders you bid half your value. With four you bid three quarters. With ten you bid nine tenths, and as the field grows the shading vanishes, because the chance that shading costs you the object rises while the saving on each pound stays the same.

Example. Values are uniform on [0,1] and there are four bidders. What should a bidder with a value of 0.8 bid, and what is their expected payoff?

The bid is 0.75×0.8=0.6. They win when all three rivals have values below 0.8, which has probability 0.83=0.512, and when they win they earn 0.8-0.6=0.2. The expected payoff is 0.2×0.512=0.1024.

Now you. Same setting with ten bidders. What does the bidder with value 0.8 bid, and what is their expected payoff now?

Answer

The bid is 0.9×0.8=0.72, and the win probability is 0.89=0.134. The expected payoff is (0.8-0.72)(0.134)=0.0107, a tenth of what it was with four bidders. Competition destroys bidder surplus twice over, by cutting the margin and by cutting the chance of collecting it, and the seller takes both.

Revenue equivalence

Two formats, two completely different bidding rules. Which raises more?

Work out the second-price revenue first. Everyone bids their value, so the price is the second-highest of n draws from the uniform distribution on [0,1]. The expected value of the kth highest of n such draws is (n+1-k)/(n+1), so the expected revenue is (n-1)/(n+1).

Now the first-price revenue. The winner is the bidder with the highest value, whose expectation is n/(n+1), and they pay (n-1)/n of it. Multiplying:

n-1n×nn+1=n-1n+1

Identical. With two bidders both formats raise 1/3 on average; with five, 2/3; with ten, 9/11. The seller is indifferent, and so, it turns out, is every bidder before learning their value.

This is the revenue equivalence theorem, proved in general by Roger Myerson and by John Riley and William Samuelson, both in 1981. Any auction mechanism in which the object always goes to the bidder with the highest value, and in which a bidder with the lowest possible value expects zero, raises the same expected revenue, whatever its rules look like. The result is a strong warning against auction design by intuition: the choice between formats cannot be justified by revenue under these assumptions, so any real argument for one has to point at an assumption that fails.

Four of them fail regularly and each names a real design decision. Risk aversion breaks the tie in favour of first-price, since a risk-averse bidder bids nearer their value to reduce the chance of losing. Correlated values break it the other way: Paul Milgrom and Robert Weber showed in 1982 that when values are affiliated, the English auction raises most and the first-price sealed bid least, because open bidding reveals information that reduces the winner's fear of overpaying. Collusion is easier in second-price and English formats, where a ring can allocate the object internally and one member bids low with no risk of being outbid by a defector who would have to pay their own bid. And entry matters more than any of it: a format that discourages weak bidders from turning up loses more revenue than any shading effect, which is why sealed bids are often preferred where one bidder is known to be strong.

The winner's curse

Everything above assumes private values: what the painting is worth to you says nothing about what it is worth to me. Change that assumption and the arithmetic turns hostile.

In a common value auction the object is worth the same to everyone, and nobody knows what that is. An offshore oil tract holds a quantity of oil that is what it is; each bidder has a geologist producing an estimate. Suppose the true value is V and each of n bidders receives a signal drawn uniformly from [V-10,V+10], unbiased in the sense that its expected value is V.

The trap is that you do not win at random. You win when your signal is the highest of the n, and the highest of n draws from that distribution has expected value

V-10+20nn+1=V+10n-1n+1

With ten bidders that is V+8.18. A bidder who treats their unbiased signal as an estimate of the value, bids close to it and wins has systematically overpaid by more than eight, and the more rivals they beat the worse it is. This is the winner's curse, named by Ed Capen, Bob Clapp and Bill Campbell in a 1971 paper in the Journal of Petroleum Technology, written after observing that returns on Gulf of Mexico leases in the 1950s and 1960s were far below what the geology had promised.

The repair is a conditioning argument, and it belongs on the list of things this course exists to teach. Do not ask what the object is worth given your signal. Ask what it is worth given your signal and given that your signal was the highest, since those are the only cases in which you pay anything. Formally, bid on E[V|si=s,si=max] rather than on E[V|si=s]. In equilibrium bidders shade for this reason on top of the ordinary first-price shading, and the equilibrium does not lose money.

Real bidders do not do this reliably. Max Bazerman and William Samuelson's 1983 experiment auctioned jars of coins worth 8toclassesofMBAstudents.Theaveragebidwas5.13, so the group as a whole was cautious. The average winning bid was 10.01,sothewinnerslostanaverageof2.01 per jar. Caution in the aggregate is no protection: the auction selects the most optimistic estimate in the room and hands them the bill.

Example. Ten bidders each receive a signal uniform within £10 million of the true value of a tract. You receive £58 million and win. What is a sensible estimate of the tract's value?

The expected overshoot of the highest of ten signals is 10×9/11=£8.18 million, so conditional on your signal being the highest, an unbiased estimate of the value is 58-8.18=£49.8 million. Your bid should be below that, not below £58 million, and the gap between those two numbers is what has bankrupted bidders in real lease sales.

Now you. The same tract and the same signal of £58 million, but only four bidders. What is the estimate now, and why has it moved?

Answer

The overshoot is 10×3/5=£6 million, so the estimate is £52 million. Fewer rivals means winning is weaker evidence that you were the optimist, so the curse is milder. The general and slightly alarming implication is that a bidder should get more cautious as the field gets larger, which is the opposite of the instinct competition produces.

What the theory has actually built

Auction theory is the part of this subject with the clearest record of practical effect, and the record includes both directions.

The New Zealand government ran a second-price sealed-bid auction for radio spectrum in 1990 on exactly the truth-telling logic above, and discovered its political weakness: the gap between winning bids and prices paid is public. One licence attracted a winning bid of NZ100,000andsoldforNZ6, and another drew NZ7millionandsoldforNZ5,000, because in each case the runner-up was far behind. The theory was working as designed and the format was abandoned.

The United Kingdom's auction of five third-generation mobile licences in April 2000, designed with advice from Ken Binmore and Paul Klemperer, ran as an ascending auction over 150 rounds and raised £22.5 billion. Germany's auction four months later raised about €50 billion. Other European countries using different rules in the same year and the same industry raised radically less, and Klemperer's account of why puts almost none of the weight on the format's revenue properties in theory and almost all of it on entry and collusion: Switzerland's auction, which allowed joint bidding until the number of bidders had fallen close to the number of licences, raised roughly €20 per head of population against the United Kingdom's €650.

That is the honest summary of what the machinery is worth. The equilibrium calculations are correct and second-order. What decides an auction is how many serious bidders show up and whether they can quietly agree not to compete, and the contribution of theory is mostly that it says clearly which of those questions to ask.

Where this leaves us

Private information turned out not to break the framework. Types, a common prior and Bayes-Nash equilibrium reproduce everything the earlier lessons did, with strategies as functions rather than actions, and they deliver sharp results: truth-telling under a second-price rule, a shading factor of (n-1)/n, revenue equivalence across formats, and a correction for the curse of winning.

The course now has its full toolkit, and one obligation left. Across fourteen lessons the theory has predicted penalty kicks accurately, predicted ultimatum offers badly, produced a folk theorem that predicts nothing, and produced auction advice that governments paid for and used. The last lesson lays those results side by side, asks what distinguishes the successes from the failures, and looks at the models built to fit the failures: level-k reasoning, quantal response and social preferences. It also confronts the temptation named in the very first lesson, that any behaviour at all can be explained after the fact by rewriting the payoffs.

Where the predictions fail

A theory earns its standing from the cases where it could have been wrong and was not, so the last thing this course owes is the record.

The scorecard

Start with the successes, because they are real and they are specific.

The zero-sum lesson's prediction about penalty kicks held: professional kickers went left about 40 per cent of the time against a predicted 38.5, keepers dived left about 42 against a predicted 42, and the sequences passed tests of serial independence that amateurs reliably fail. Mark Walker and John Wooders found the same in the serve directions of top tennis players in 2001. Where the game is genuinely zero-sum, the stakes are high, the players are professionals and the choice is made thousands of times, minimax is an accurate description of behaviour.

Auction theory has been used to design real auctions for real money, and its predictions about how bidders shade, how the winner's curse bites and how entry and collusion matter have survived contact with billions of pounds of spectrum. Matching theory, the branch of the subject not covered here, redesigned the American medical residency match in 1998, the New England kidney exchange from 2004, and the school assignment systems of New York and Boston, in each case by finding the strategic flaw in an existing procedure and removing it.

Now the failures, stated as sharply.

Sole defection in a one-shot prisoner's dilemma is predicted for everyone and observed in about half to two thirds of first plays, with cooperation persisting at some rate however much experience subjects get. In finitely repeated prisoner's dilemmas, backward induction predicts defection from round one and subjects cooperate for most of the sequence, defecting only near the end. In the centipede game, backward induction says the first player takes immediately; in Richard McKelvey and Thomas Palfrey's 1992 experiments, fewer than one pair in ten did. In public goods games, the dominant strategy is to contribute nothing, and subjects typically contribute around half their endowment in the first round, declining but not vanishing over ten rounds. And in the ultimatum game the prediction fails at both ends, with proposers offering near half and responders refusing money.

The pattern is not random. Predictions succeed where interests are strictly opposed, where the same situation recurs often enough to be learned, and where the stakes are large enough to buy attention. They fail where a cooperative or fair outcome is available and salient, and where the reasoning required runs to several steps of "they know that I know". Both halves of that pattern have been modelled.

The beauty contest

The cleanest experiment in the subject isolates the second failure. Everybody picks a number between 0 and 100, and whoever comes closest to two thirds of the group average wins. John Maynard Keynes used the image in 1936, comparing investors to entrants in a newspaper competition to pick the faces other entrants would pick; Rosemarie Nagel turned it into a laboratory game in 1995.

The equilibrium is easy. If everyone picks below 100, two thirds of the average is below 67, so nothing above 67 can win, so nobody should pick above 67. But then nothing above 44 can win, and so on down. The only Nash equilibrium is everybody choosing 0, and it survives iterated deletion of dominated strategies, which is the lesson-two argument applied about fifty times over.

Nobody plays 0. Averages in the laboratory come out between 20 and 40, and a famous version run by the Financial Times in 1997 drew a mean guess of 18.9 with a winning number of 13. The distribution of choices is not scattered noise either: it clusters at recognisable points.

Those points are what a chain of reasoning produces if it stops early. Suppose a naive player picks at random, averaging 50. Someone who best responds to that picks 23(50)=33.3. Someone who best responds to them picks 22.2, then 14.8, then 9.9. The observed spikes in the data sit at 33 and 22, which is exactly what you see if most people do one or two rounds of the reasoning and stop.

Example. The same game with a target of 0.7 times the average. What do players doing zero, one, two and three rounds of reasoning choose?

A zero-round player picks at random, averaging 50. One round gives 0.7(50)=35. Two rounds give 0.7(35)=24.5. Three rounds give 0.7(24.5)=17.15. The equilibrium is still 0, and the practical prediction, that the winning number will be somewhere in the twenties, comes from the depth of reasoning rather than from the equilibrium.

Now you. With the two thirds target, what would a group of players who all reason to exactly two rounds produce as the winning number, and what would happen if the same group played again?

Answer

They would all pick 22.2, so the average is 22.2 and two thirds of it is 14.8, meaning the winners are whoever went one round further. Played again, everyone anchors on the observed 22.2 rather than on 50, and the choices drop to around 14.8, then 9.9. This is exactly what happens in repeated sessions: guesses fall towards zero over rounds, which shows the equilibrium is being learned rather than deduced.

Level-k and cognitive hierarchy

Turn that observation into a model. Level-k theory assumes a population of types. A level-0 player is non-strategic, choosing randomly or by some salient rule. A level-1 player best responds to level-0. A level-2 player best responds to level-1, and so on. No player believes anyone is smarter than themselves, which is precisely the assumption that common knowledge of rationality forbids.

Fitting it to beauty contest data puts most subjects at level 1 or 2 with almost nobody above 3. Colin Camerer, Teck-Hua Ho and Juin-Kuan Chong's 2004 cognitive hierarchy variant makes a level-k player best respond to a mixture of all lower levels rather than to level k-1 alone, with the levels distributed Poisson; across many games the fitted mean is around 1.5.

What makes this more than curve fitting is that it predicts across games. The same fitted distribution of levels explains why people play close to equilibrium in games where equilibrium requires one step of reasoning and far from it where it requires many, and it predicts which games those are before seeing the data. It also delivers a practical warning: an equilibrium reached by a long chain of iterated deletion, like the chain store paradox of the credibility lesson, should be trusted much less than one reached in a single step. That was flagged in the dominance lesson on theoretical grounds and is now an empirical fact.

Mistakes in proportion to their cost

The second repair keeps equilibrium and drops perfection. In quantal response equilibrium, introduced by McKelvey and Palfrey in 1995, players do not always choose the best action; they choose better actions more often, with the choice probabilities given by a logit rule

P(a)=eλu(a)aeλu(a)

and beliefs that are correct about these noisy choices, which is what keeps it an equilibrium concept rather than a description of confusion.

The parameter λ measures precision. At λ=0 play is uniformly random. As λ the best action is taken with probability one and quantal response equilibrium becomes Nash equilibrium. In between, the model says something Nash cannot: that mistakes are frequent when they are cheap and rare when they are expensive.

That single idea resolves several anomalies at once. It explains why subjects deviate more in games where the payoff differences are small, why cooperation in the finitely repeated prisoner's dilemma survives many rounds and collapses at the end, where the cost of cooperating stops being offset by anything, and why passing in the centipede game is common early, where a mistake costs little, and rare at the last node.

Example. Two actions pay 5 and 4. What is the probability of choosing the better one at λ=1, and at λ=3?

The logit probability is eλ5/(eλ5+eλ4), which simplifies to 1/(1+e-λ) because only the payoff difference of 1 matters. At λ=1 this is 1/(1+e-1)=0.731, and at λ=3 it is 1/(1+e-3)=0.953. A threefold rise in precision moves the error rate from 27 per cent to 5 per cent.

Now you. The same two actions but with payoffs 50 and 40 rather than 5 and 4, at λ=1. What is the probability now, and what does that say about scaling payoffs?

Answer

The difference is now 10, so the probability is 1/(1+e-10)=0.99995. Multiplying every payoff by ten has made errors essentially disappear. That is the model's most important property and its most awkward one: unlike Nash equilibrium, quantal response is not invariant to the affine rescaling of payoffs that the first lesson established as harmless, so λ is not a pure measure of rationality but is tied to the units the game is written in.

Preferences that are not selfish

The third repair keeps rationality intact and changes what people want. Ultimatum responders refusing money are behaving inconsistently only if their payoff is the money, and the first lesson was explicit that a payoff is a utility rather than a cash amount.

Ernst Fehr and Klaus Schmidt's 1999 model of inequity aversion writes player i's utility in a two-person allocation as

Ui=xi-αimax(xj-xi,0)-βimax(xi-xj,0)

with αiβi0. The first penalty is envy, the discomfort of getting less than the other; the second is guilt, weaker, from getting more.

Apply it to an ultimatum responder offered y out of 10. Accepting gives y-α(10-2y) when y<5; rejecting gives 0, since both then have nothing and there is no inequity. So the offer is accepted when y(1+2α)10α, that is

y10α1+2α

A responder with α=0.5 rejects anything below £2.50; one with α=1 rejects below £3.33; a purely selfish responder with α=0 accepts anything. Fehr and Schmidt fitted a distribution of α across the population and used the same distribution, unchanged, to predict behaviour in public goods games, market games and gift exchange. In competitive market experiments where one side is rationed, the same fair-minded subjects produce outcomes indistinguishable from the selfish prediction, because competition removes the ability to act on the preference. Predicting fairness in one institution and its complete absence in another, from one set of parameters, is what a good behavioural model looks like.

Example. A responder with α=1.2 faces a £10 pie. What is the smallest offer they accept?

y10(1.2)/(1+2.4)=12/3.4=£3.53. Fehr and Schmidt's fitted population has about 30 per cent of subjects at α=0, 30 per cent at 0.5, 30 per cent at 1 and 10 per cent at 4, which produces a rejection rate rising steeply below about a third of the pie, matching the data.

Now you. Using that fitted population, roughly what fraction of responders reject an offer of £2 out of £10?

Answer

The thresholds are £0 for α=0, £2.50 for α=0.5, £3.33 for α=1 and £4.44 for α=4. An offer of £2 falls below every threshold except the first, so the 70 per cent of the population with α>0 reject it. Observed rejection rates for offers around a fifth of the pie are lower than that, nearer a half, which is a reminder that the fitted parameters are a summary rather than a measurement of anybody.

The escape hatch, and the discipline

Three repairs, each successful, and each raising the objection the first lesson promised to return to. If a prediction fails, the payoffs can always be rewritten until it succeeds. Cooperation in the prisoner's dilemma? The players must value each other's welfare. Rejection of a fair-looking offer? Inequity aversion. Anything at all? Some utility function makes it optimal.

Taken to its end this makes the theory unfalsifiable, and the fact that it can be done is not a defence. What separates the good work from the bad is not whether a preference parameter was added but whether it was disciplined afterwards, and there are three standard disciplines.

Fit here, predict there. A parameter estimated on ultimatum games and then used, unchanged, to predict public goods contributions is doing scientific work. One estimated separately for each game is a description with extra steps.

Count the parameters. Quantal response equilibrium adds one number, λ, to the whole framework. Cognitive hierarchy adds one, τ. Inequity aversion adds two per player and a population distribution. Each buys a great deal of explanatory power for very little, and a model that adds a parameter per anomaly buys nothing.

Test the process, not just the choice. If a subject is playing level-2, they should look up the payoffs a level-2 calculation needs and not others. Experiments recording which payoff cells subjects actually open, and how long they take, have been used to distinguish level-k from quantal response in cases where both fit the choices equally well. Predictions about the reasoning rather than the outcome are the hardest kind to fudge.

The corresponding discipline for applied work is blunter. If a model of a situation reproduces the observed behaviour only after the payoffs have been adjusted to fit it, it has not explained the behaviour, and its predictions about a changed situation are worthless.

What you can do now

The subject that started with two airlines and a fare has covered a specific set of skills, and it is worth naming them as a checklist for any strategic situation.

Write it down: players, strategy sets, payoffs, and whether the payoffs are utilities or a convenient stand-in. Strip it with dominance, remembering that the deeper rounds of iterated deletion are the least reliable part of the analysis. Find the equilibria, in mixed strategies where none exists in pure ones, and count them honestly rather than reporting the one you like. If the choices are continuous, solve the best-response functions. If moves are sequenced, build the tree and use backward induction, then ask which of the surviving threats anybody would actually carry out. If the situation recurs, compute the patience threshold, and remember that repetition sustains bad arrangements as readily as good ones. If the question is how to split a surplus, take the outside options off the top and look at who can afford to wait. If the parties know things about themselves that others do not, model the types and the beliefs, and ask what winning would tell you.

Then ask the question this lesson exists for. Is this a situation of the kind where the theory has a good record, meaning opposed interests, repeated play, experienced participants and real stakes? Or is it one where it has a poor one, meaning a single encounter, a fairness norm in view, or a conclusion resting on ten steps of mutual reasoning? The models in this lesson give the second case a first correction rather than a shrug.

The framework does not tell you what will happen. It tells you what a situation rewards, which pressures are present, and which of your beliefs about the other side your conclusion is resting on. That last one is the most useful, and it is available from nothing else.

Game Theory, from libre.university