A theory earns its standing from the cases where it could have been wrong and was not, so the last thing this course owes is the record.
The scorecard
Start with the successes, because they are real and they are specific.
The zero-sum lesson's prediction about penalty kicks held: professional kickers went left about 40 per cent of the time against a predicted 38.5, keepers dived left about 42 against a predicted 42, and the sequences passed tests of serial independence that amateurs reliably fail. Mark Walker and John Wooders found the same in the serve directions of top tennis players in 2001. Where the game is genuinely zero-sum, the stakes are high, the players are professionals and the choice is made thousands of times, minimax is an accurate description of behaviour.
Auction theory has been used to design real auctions for real money, and its predictions about how bidders shade, how the winner's curse bites and how entry and collusion matter have survived contact with billions of pounds of spectrum. Matching theory, the branch of the subject not covered here, redesigned the American medical residency match in 1998, the New England kidney exchange from 2004, and the school assignment systems of New York and Boston, in each case by finding the strategic flaw in an existing procedure and removing it.
Now the failures, stated as sharply.
Sole defection in a one-shot prisoner's dilemma is predicted for everyone and observed in about half to two thirds of first plays, with cooperation persisting at some rate however much experience subjects get. In finitely repeated prisoner's dilemmas, backward induction predicts defection from round one and subjects cooperate for most of the sequence, defecting only near the end. In the centipede game, backward induction says the first player takes immediately; in Richard McKelvey and Thomas Palfrey's 1992 experiments, fewer than one pair in ten did. In public goods games, the dominant strategy is to contribute nothing, and subjects typically contribute around half their endowment in the first round, declining but not vanishing over ten rounds. And in the ultimatum game the prediction fails at both ends, with proposers offering near half and responders refusing money.
The pattern is not random. Predictions succeed where interests are strictly opposed, where the same situation recurs often enough to be learned, and where the stakes are large enough to buy attention. They fail where a cooperative or fair outcome is available and salient, and where the reasoning required runs to several steps of "they know that I know". Both halves of that pattern have been modelled.
The beauty contest
The cleanest experiment in the subject isolates the second failure. Everybody picks a number between 0 and 100, and whoever comes closest to two thirds of the group average wins. John Maynard Keynes used the image in 1936, comparing investors to entrants in a newspaper competition to pick the faces other entrants would pick; Rosemarie Nagel turned it into a laboratory game in 1995.
The equilibrium is easy. If everyone picks below 100, two thirds of the average is below 67, so nothing above 67 can win, so nobody should pick above 67. But then nothing above 44 can win, and so on down. The only Nash equilibrium is everybody choosing 0, and it survives iterated deletion of dominated strategies, which is the lesson-two argument applied about fifty times over.
Nobody plays 0. Averages in the laboratory come out between 20 and 40, and a famous version run by the Financial Times in 1997 drew a mean guess of 18.9 with a winning number of 13. The distribution of choices is not scattered noise either: it clusters at recognisable points.
Those points are what a chain of reasoning produces if it stops early. Suppose a naive player picks at random, averaging 50. Someone who best responds to that picks . Someone who best responds to them picks 22.2, then 14.8, then 9.9. The observed spikes in the data sit at 33 and 22, which is exactly what you see if most people do one or two rounds of the reasoning and stop.
Example. The same game with a target of 0.7 times the average. What do players doing zero, one, two and three rounds of reasoning choose?
A zero-round player picks at random, averaging 50. One round gives . Two rounds give . Three rounds give . The equilibrium is still 0, and the practical prediction, that the winning number will be somewhere in the twenties, comes from the depth of reasoning rather than from the equilibrium.
Now you. With the two thirds target, what would a group of players who all reason to exactly two rounds produce as the winning number, and what would happen if the same group played again?
Answer
They would all pick 22.2, so the average is 22.2 and two thirds of it is 14.8, meaning the winners are whoever went one round further. Played again, everyone anchors on the observed 22.2 rather than on 50, and the choices drop to around 14.8, then 9.9. This is exactly what happens in repeated sessions: guesses fall towards zero over rounds, which shows the equilibrium is being learned rather than deduced.
Level-k and cognitive hierarchy
Turn that observation into a model. Level-k theory assumes a population of types. A level-0 player is non-strategic, choosing randomly or by some salient rule. A level-1 player best responds to level-0. A level-2 player best responds to level-1, and so on. No player believes anyone is smarter than themselves, which is precisely the assumption that common knowledge of rationality forbids.
Fitting it to beauty contest data puts most subjects at level 1 or 2 with almost nobody above 3. Colin Camerer, Teck-Hua Ho and Juin-Kuan Chong's 2004 cognitive hierarchy variant makes a level- player best respond to a mixture of all lower levels rather than to level alone, with the levels distributed Poisson; across many games the fitted mean is around 1.5.
What makes this more than curve fitting is that it predicts across games. The same fitted distribution of levels explains why people play close to equilibrium in games where equilibrium requires one step of reasoning and far from it where it requires many, and it predicts which games those are before seeing the data. It also delivers a practical warning: an equilibrium reached by a long chain of iterated deletion, like the chain store paradox of the credibility lesson, should be trusted much less than one reached in a single step. That was flagged in the dominance lesson on theoretical grounds and is now an empirical fact.
Mistakes in proportion to their cost
The second repair keeps equilibrium and drops perfection. In quantal response equilibrium, introduced by McKelvey and Palfrey in 1995, players do not always choose the best action; they choose better actions more often, with the choice probabilities given by a logit rule
and beliefs that are correct about these noisy choices, which is what keeps it an equilibrium concept rather than a description of confusion.
The parameter measures precision. At play is uniformly random. As the best action is taken with probability one and quantal response equilibrium becomes Nash equilibrium. In between, the model says something Nash cannot: that mistakes are frequent when they are cheap and rare when they are expensive.
That single idea resolves several anomalies at once. It explains why subjects deviate more in games where the payoff differences are small, why cooperation in the finitely repeated prisoner's dilemma survives many rounds and collapses at the end, where the cost of cooperating stops being offset by anything, and why passing in the centipede game is common early, where a mistake costs little, and rare at the last node.
Example. Two actions pay 5 and 4. What is the probability of choosing the better one at , and at ?
The logit probability is , which simplifies to because only the payoff difference of 1 matters. At this is , and at it is . A threefold rise in precision moves the error rate from 27 per cent to 5 per cent.
Now you. The same two actions but with payoffs 50 and 40 rather than 5 and 4, at . What is the probability now, and what does that say about scaling payoffs?
Answer
The difference is now 10, so the probability is . Multiplying every payoff by ten has made errors essentially disappear. That is the model's most important property and its most awkward one: unlike Nash equilibrium, quantal response is not invariant to the affine rescaling of payoffs that the first lesson established as harmless, so is not a pure measure of rationality but is tied to the units the game is written in.
Preferences that are not selfish
The third repair keeps rationality intact and changes what people want. Ultimatum responders refusing money are behaving inconsistently only if their payoff is the money, and the first lesson was explicit that a payoff is a utility rather than a cash amount.
Ernst Fehr and Klaus Schmidt's 1999 model of inequity aversion writes player 's utility in a two-person allocation as
with . The first penalty is envy, the discomfort of getting less than the other; the second is guilt, weaker, from getting more.
Apply it to an ultimatum responder offered out of 10. Accepting gives when ; rejecting gives 0, since both then have nothing and there is no inequity. So the offer is accepted when , that is
A responder with rejects anything below £2.50; one with rejects below £3.33; a purely selfish responder with accepts anything. Fehr and Schmidt fitted a distribution of across the population and used the same distribution, unchanged, to predict behaviour in public goods games, market games and gift exchange. In competitive market experiments where one side is rationed, the same fair-minded subjects produce outcomes indistinguishable from the selfish prediction, because competition removes the ability to act on the preference. Predicting fairness in one institution and its complete absence in another, from one set of parameters, is what a good behavioural model looks like.
Example. A responder with faces a £10 pie. What is the smallest offer they accept?
. Fehr and Schmidt's fitted population has about 30 per cent of subjects at , 30 per cent at 0.5, 30 per cent at 1 and 10 per cent at 4, which produces a rejection rate rising steeply below about a third of the pie, matching the data.
Now you. Using that fitted population, roughly what fraction of responders reject an offer of £2 out of £10?
Answer
The thresholds are £0 for , £2.50 for , £3.33 for and £4.44 for . An offer of £2 falls below every threshold except the first, so the 70 per cent of the population with reject it. Observed rejection rates for offers around a fifth of the pie are lower than that, nearer a half, which is a reminder that the fitted parameters are a summary rather than a measurement of anybody.
The escape hatch, and the discipline
Three repairs, each successful, and each raising the objection the first lesson promised to return to. If a prediction fails, the payoffs can always be rewritten until it succeeds. Cooperation in the prisoner's dilemma? The players must value each other's welfare. Rejection of a fair-looking offer? Inequity aversion. Anything at all? Some utility function makes it optimal.
Taken to its end this makes the theory unfalsifiable, and the fact that it can be done is not a defence. What separates the good work from the bad is not whether a preference parameter was added but whether it was disciplined afterwards, and there are three standard disciplines.
Fit here, predict there. A parameter estimated on ultimatum games and then used, unchanged, to predict public goods contributions is doing scientific work. One estimated separately for each game is a description with extra steps.
Count the parameters. Quantal response equilibrium adds one number, , to the whole framework. Cognitive hierarchy adds one, . Inequity aversion adds two per player and a population distribution. Each buys a great deal of explanatory power for very little, and a model that adds a parameter per anomaly buys nothing.
Test the process, not just the choice. If a subject is playing level-2, they should look up the payoffs a level-2 calculation needs and not others. Experiments recording which payoff cells subjects actually open, and how long they take, have been used to distinguish level-k from quantal response in cases where both fit the choices equally well. Predictions about the reasoning rather than the outcome are the hardest kind to fudge.
The corresponding discipline for applied work is blunter. If a model of a situation reproduces the observed behaviour only after the payoffs have been adjusted to fit it, it has not explained the behaviour, and its predictions about a changed situation are worthless.
What you can do now
The subject that started with two airlines and a fare has covered a specific set of skills, and it is worth naming them as a checklist for any strategic situation.
Write it down: players, strategy sets, payoffs, and whether the payoffs are utilities or a convenient stand-in. Strip it with dominance, remembering that the deeper rounds of iterated deletion are the least reliable part of the analysis. Find the equilibria, in mixed strategies where none exists in pure ones, and count them honestly rather than reporting the one you like. If the choices are continuous, solve the best-response functions. If moves are sequenced, build the tree and use backward induction, then ask which of the surviving threats anybody would actually carry out. If the situation recurs, compute the patience threshold, and remember that repetition sustains bad arrangements as readily as good ones. If the question is how to split a surplus, take the outside options off the top and look at who can afford to wait. If the parties know things about themselves that others do not, model the types and the beliefs, and ask what winning would tell you.
Then ask the question this lesson exists for. Is this a situation of the kind where the theory has a good record, meaning opposed interests, repeated play, experienced participants and real stakes? Or is it one where it has a poor one, meaning a single encounter, a fairness norm in view, or a conclusion resting on ten steps of mutual reasoning? The models in this lesson give the second case a first correction rather than a shrug.
The framework does not tell you what will happen. It tells you what a situation rewards, which pressures are present, and which of your beliefs about the other side your conclusion is resting on. That last one is the most useful, and it is available from nothing else.