Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

What a game is

A decision problem becomes a game the moment the best thing for you to do depends on what somebody else decides at the same time, and the whole subject exists because that dependence breaks ordinary decision-making.

Against nature, against a person

Deciding whether to carry an umbrella is a problem with one decision-maker. The weather has no interest in whether you get wet, so you can attach a probability to rain, weigh the cost of carrying the umbrella against the cost of a soaking, and take the option with the better expected payoff. That is decision theory, and the machinery for it is the probability course: a distribution over states of the world, a payoff for each combination of state and action, and an expectation to maximise.

Now change one thing. Suppose the rain is produced by somebody who is paid whenever you get wet, and who knows you are reasoning about umbrellas. There is no fixed probability of rain to look up, because the chance of rain depends on the chance you carry an umbrella, which depends on what you think the chance of rain is. Each side's calculation is an input to the other's. The regress is real, not a trick of presentation, and no amount of care about probabilities dissolves it.

Game theory is the machinery for that second case. It was assembled for the first time in John von Neumann and Oskar Morgenstern's Theory of Games and Economic Behavior in 1944, on foundations von Neumann had laid in a 1928 paper on parlour games, and the word "game" has been misleading people ever since. The subject is applied to price wars, arms control, evolution, auctions, plea bargains, divorce settlements and the timing of an election. What makes something a game is not that it is played for fun but that the payoff to your choice is not yours alone to determine.

The three ingredients

Strip the story away and every strategic situation has the same skeleton, and it has exactly three parts.

There is a set of players, indexed i=1,2,,n, meaning the decision-makers whose choices matter. There is, for each player, a set of strategies Si, meaning the options that player can take. And there is, for each player, a payoff function ui that assigns a number to every complete combination of choices. A combination with one strategy for each player is a strategy profile, written s=(s1,s2,,sn), and the whole point of the payoff function is that it takes the entire profile as its argument: ui(s), not ui(si). There is a standard shorthand for what everybody except player i is doing, s-i, so a payoff is written ui(si,s-i) when the emphasis is on separating your own choice from the rest.

Those three ingredients together are the game in normal form, sometimes called strategic form. A two-player game with a short strategy list each is drawn as a grid: one player picks the row, the other picks the column, and the cell where they meet holds two numbers, the row player's payoff first and the column player's second. The convention on the order of those two numbers is universal and worth committing to memory, because reading a grid with them the wrong way round produces confident nonsense.

Here is one. Two airlines fly the same route, and each sets a fare of either £120 or £90 for the season. The route carries 200 passengers whatever the fares, and it costs an airline £30 to carry one. At equal fares the traffic splits evenly. If one undercuts, it takes 170 passengers and leaves 30 to its rival. Multiplying out gives four cells, in thousands of pounds:

A's fareB charges £120B charges £90
£1209.0, 9.02.7, 10.2
£9010.2, 2.76.0, 6.0

Every number there was computed, not invented. If both hold at £120 they carry 100 passengers each at a margin of £90, which is £9,000. If A alone cuts to £90, its margin falls to £60 but it carries 170, which is £10,200, while B is left with 30 passengers at a £90 margin, which is £2,700. The table is doing real work: read across A's top row and you see A's fortune swinging between £9,000 and £2,700 on a decision that belongs to B.

Example. Suppose loyalty is weaker than assumed and the undercutter takes 185 of the 200 passengers rather than 170. What do the two off-diagonal cells become?

The margins are unchanged, £90 at the high fare and £60 at the low one, so only the passenger counts move. The undercutter earns 185×60=11{,}100 and the rival earns 15×90=1{,}350. The off-diagonal cells become 11.1, 1.35 and 1.35, 11.1 in thousands. Undercutting became more attractive and being undercut became worse, while the two diagonal cells did not move at all.

Now you. Return to the original 170/30 split, but suppose the cost of carrying a passenger rises from £30 to £50. What are the four cells now?

Answer

The margins become £70 at the high fare and £40 at the low one. Both high: 100×70=7{,}000 each. Both low: 100×40=4{,}000 each. Undercutter: 170×40=6{,}800; rival: 30×70=2{,}100. So the grid reads 7.0, 7.0 across the top left, 2.1, 6.8 and 6.8, 2.1 off the diagonal, and 4.0, 4.0 at the bottom right, in thousands of pounds.

Payoffs are utilities, not money

The numbers in the grid look like money, and in the airline example they are. That is a convenience, and taking it as the definition causes trouble as soon as the analysis involves any uncertainty at all, which from the next lesson but one it always will.

What a payoff has to represent is preference over lotteries, meaning probability distributions over outcomes. That is a stronger requirement than ranking the outcomes themselves. Von Neumann and Morgenstern's contribution alongside the games was a representation theorem: if a person's preferences over lotteries satisfy four conditions, roughly that any two lotteries can be compared, that comparisons are transitive, that a preference is not reversed by mixing both sides with a third lottery, and that there are no infinitely good or infinitely bad outcomes, then there exists a function u on outcomes such that the person prefers one lottery to another exactly when it has the higher expected value of u. The payoff numbers in a game are that u.

The practical consequence is that payoffs already contain the player's attitude to risk, so the analysis never needs to add it afterwards. Take a person whose utility for a sum of money m is u(m)=m. A coin flip between nothing and £100 has expected utility 0.50+0.5100=5, and since 25=5, that gamble is worth exactly £25 to them, against an expected money value of £50. They would rather have a certain £36, worth 36=6, than the coin flip, even though the coin flip pays more on average. Put £36 and £50 in a payoff table and you would predict the wrong choice; put 6 and 5 in it and you predict the right one.

Example. The same person, with u(m)=m, is offered a coin flip between £49 and £121. What certain sum is that gamble worth to them, and how does it compare with its average payout?

Expected utility is 0.549+0.5121=0.5(7)+0.5(11)=9. The certain sum with that utility is the one whose square root is 9, so £81. The gamble pays £85 on average, so the person would trade it for £81 in cash, giving up £4 of expected money to be rid of the risk. That £4 is the risk premium, and it is what the curvature of m encodes.

Now you. Same utility function. A gamble pays £16 with probability 0.5 and £144 with probability 0.5. What is its certainty equivalent?

Answer

0.516+0.5144=0.5(4)+0.5(12)=8, and 82=64, so the gamble is worth a certain £64. Its average payout is £80, so the risk premium is £16.

There is a converse warning. Because payoffs are utilities rather than money, a player who maximises their payoff is not thereby selfish. If you would genuinely rather split a windfall than keep it, that preference belongs in your utility numbers, and a model in which everyone maximises ui is a model in which everyone pursues what they actually want. This is a real strength of the framework and also its most abused escape hatch, because any observed behaviour whatsoever can be rationalised after the fact by adjusting the payoffs. The last lesson of this course is largely about that temptation.

What can be changed without changing the game

A utility function is not unique. If u represents someone's preferences over lotteries, so does v=αu+β for any constant α>0 and any constant β, because expectations are linear: E[v]=αE[u]+β preserves every comparison. This is called invariance under positive affine transformation, and it is the exact analogue of measuring temperature in Celsius or Fahrenheit.

So the same game can be written with different numbers. A player's payoffs may be doubled, or have 50 added, without changing a single strategic conclusion. What is not allowed is applying a transformation that is not affine, such as squaring, and what is emphatically not allowed is comparing one player's payoff with another's. A cell reading 9.0, 9.0 does not mean the players are equally well off, and a cell reading 10.2, 2.7 does not mean the first player gains four times what the second suffers. Interpersonal comparison needs an extra assumption, and the bargaining lessons later in this course will have to make one explicitly.

Example. Player 1's payoffs across four cells are 0, 1, 3, 4. A colleague writes the same game with 2, 5, 11, 14. Are these the same preferences?

Test whether one constant multiplier and one constant offset carry all four across. From 0 to 2 the offset is 2 if the multiplier is anything. From 1 to 5: α(1)+2=5 gives α=3. Check the rest: 3(3)+2=11 and 3(4)+2=14, both right. It is the transformation v=3u+2 with α=3>0, so the two tables describe the same player.

Now you. Player 2's payoffs are 1, 2, 4 and a colleague writes 1, 4, 16. Same preferences?

Answer

No. Matching the first two needs α(1)+β=1 and α(2)+β=4, so α=3 and β=-2. That predicts 3(4)-2=10 for the third, not 16. The colleague squared the payoffs, which is not an affine transformation, and it changes how the player ranks lotteries: a coin flip between 1 and 16 beats a certain 4 in the second table and loses to it in the first.

Actions, strategies, and the difference

In the airline grid a strategy is just a choice, and the words "action" and "strategy" can be used interchangeably. That stops being true the moment players move in sequence and can see what has happened.

A strategy is a complete contingent plan: it specifies what the player does at every point where they might have to act, including points that will never be reached if the plan is followed. A chess strategy in this sense is not an opening preference but a full specification of a reply to every legal position, which is why the strategy sets of real games are astronomically large and why the normal form is a theoretical device rather than something anybody writes out. The definition looks pedantic and it earns its keep later: it is what allows a sequential game to be analysed with the same equilibrium concept as a simultaneous one, and it is what makes the phrase "a threat nobody would carry out" precise.

Simultaneity matters less than it sounds, too. What the normal form assumes is not that the players move at the same instant but that neither learns the other's choice before committing. Sealed bids opened a week apart are simultaneous in the only sense the model cares about.

What the model assumes

Three assumptions are doing the work, and they should be visible rather than smuggled.

Players are rational: each has preferences satisfying the utility axioms and chooses to maximise expected payoff given their beliefs. This does not claim that people are calculating machines; it is a baseline that says what a situation rewards, and the interesting research programme of the past forty years is the catalogue of situations where the baseline misses.

The structure is common knowledge: every player knows the players, strategies and payoffs, every player knows that every player knows, and so on without limit. This is much stronger than everyone happening to know it, and the infinite regress is not decorative. Several of the arguments ahead, iterated deletion in particular, use the higher levels explicitly.

And players do not communicate or bind themselves unless the game says so. A promise, a contract or a hostage changes the game rather than the reasoning about it, which is why the lesson on commitment is about redesigning the tree rather than about being trustworthy.

Where this leaves us

The airline grid is now fully specified, and specifying it has predicted precisely nothing. Both firms would rather be in the 9.0, 9.0 cell than the 6.0, 6.0 cell, and yet neither has any way of getting there by choosing well. What is missing is a rule that takes a grid and returns a prediction.

The next lesson supplies the weakest such rule, one that needs no assumption about what your rival believes, only that they will not play something that is worse for them no matter what happens. Applied to the airline grid it gives a single answer, and the answer is the cell both firms like least.