Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Mixed strategies

A game in which one player wants to be predicted wrongly has no equilibrium among strategies that name a single action, because naming a single action is exactly what makes a player predictable.

Strategies that are lotteries

The previous lesson left the tax authority and the taxpayer chasing each other round a cycle of best responses, with no cell where both were content. Extend what a strategy is allowed to be and the cycle closes.

A mixed strategy for player i is a probability distribution σi over that player's pure strategies. In a two-action game it is one number, the probability of the first action. The original pure strategies are still available as the degenerate mixtures that put probability 1 on one action, so nothing is lost by the extension. The set of pure strategies that a mixed strategy plays with positive probability is called its support.

Payoffs extend by expectation, which is exactly what the utility axioms of the first lesson licensed. If player 1 plays Top with probability p and player 2 plays Left with probability q, then each cell of a two-by-two grid occurs with the product of its two probabilities, and the expected payoff is the weighted sum of the four cell payoffs. Expected payoff is linear in each player's own probabilities separately, and that linearity does all the work in what follows.

The definition of equilibrium does not change at all. A profile of mixed strategies is a Nash equilibrium when no player can raise their expected payoff by switching to any other strategy, mixed or pure, holding the others fixed.

The indifference condition

Linearity has a consequence that turns equilibrium from a search into an equation.

Suppose player 1's mixed strategy is a best response to what player 2 is doing, and suppose it puts positive probability on two pure strategies, Top and Bottom. Player 1's expected payoff is then a weighted average of the expected payoff of Top and the expected payoff of Bottom. If Top were worth strictly more, shifting probability towards Top would raise the average, so the current mixture would not be a best response. Therefore:

Every pure strategy in the support of a best response yields the same expected payoff, and that payoff is at least as large as any pure strategy outside the support.

The consequence is odd on first meeting and central to everything after it. In equilibrium a mixing player is exactly indifferent between the things they are mixing over. They are not randomising because randomising is better; they are randomising because it makes no difference to them, and because it makes a great deal of difference to the other player. Which leads directly to the mechanical rule: each player's probabilities are pinned down by the other player's indifference, not by their own.

Take the audit game, with the taxpayer's payoff first and figures in thousands:

TaxpayerAuditDo not audit
Evade-100, 8050, -50
Declare0, -200, 0

Let q be the probability the authority audits. The taxpayer's expected payoff from evading is -100q+50(1-q)=50-150q, and from declaring it is 0 whatever happens. Setting them equal gives q=1/3. Now let p be the probability the taxpayer evades. The authority's expected payoff from auditing is 80p-20(1-p)=100p-20, and from not auditing it is -50p. Setting those equal gives 150p=20, so p=2/150.133.

The equilibrium is that the taxpayer evades about 13.3 per cent of the time and the authority audits a third of returns. Check it: at q=1/3 the taxpayer earns 0 from either action, so any mixture including 0.133 is a best response, and at p=2/15 the authority earns -50(2/15)=-6.67 from either action, so its mixture is a best response too. Both conditions hold, so this is an equilibrium. It is the only one: the previous lesson showed the game has no pure equilibrium, so both players must be mixing, and the two indifference equations then have exactly one solution each.

Example. A cycling team chooses whether to dope and an anti-doping agency chooses whether to test, with the team's payoff first:

TeamTestNo test
Dope-70, 4030, -30
Ride clean0, -100, 0

Find the mixed equilibrium.

Neither cell of any column is stable, so both players mix. Let q be the probability of a test. The team's payoff from doping is -70q+30(1-q)=30-100q and from riding clean it is 0, so q=0.3. Let p be the probability of doping. The agency's payoff from testing is 40p-10(1-p)=50p-10 and from not testing it is -30p, so 80p=10 and p=0.125. The team dopes an eighth of the time, the agency tests three races in ten, and each side is indifferent given the other.

Now you. Find the mixed equilibrium of this game, with player 1's payoff first.

Player 1LR
T5, 20, 4
B2, 51, 2
Answer

Let q be the probability of L. Player 1 earns 5q from T and 2q+1(1-q)=1+q from B, equal when q=0.25. Let p be the probability of T. Player 2 earns 2p+5(1-p)=5-3p from L and 4p+2(1-p)=2+2p from R, equal when 5p=3, so p=0.6. The equilibrium payoffs are 1.25 to player 1 and 3.2 to player 2, and each can be checked twice by computing it from either of that player's two actions.

The comparative static nobody expects

Now use the model for what models are for. The obvious lever a government reaches for is a harsher penalty. Suppose evasion caught in an audit costs the taxpayer 300 rather than 100, because the extra severity is a criminal record rather than money, so it costs the taxpayer without adding anything to the authority's side of the table.

Recompute. The taxpayer's indifference now reads -300q+50(1-q)=0, giving q=50/350=1/70.143. The authority's indifference is untouched, because none of its own payoffs changed, so p is still 2/15.

The evasion rate does not move. What moves is the audit rate, which falls from 33.3 per cent to 14.3 per cent. Tripling the penalty bought no reduction in evasion whatsoever; it bought a cheaper enforcement budget, and the authority's expected payoff, -50p, is unchanged at -6.67 as well. If you want less evasion, the lever that works is on the authority's side of the table: make auditing cheaper or more accurate, and p falls.

This is the general shape of mixed equilibrium comparative statics, and it is deeply counterintuitive until the indifference condition is internalised. Your own payoffs determine the other player's behaviour; the other player's payoffs determine yours. The result generalises well beyond tax: in inspection games, doping controls, fare evasion and audit-like enforcement generally, raising the punishment reduces the amount of enforcement rather than the amount of the offence, unless the enforcer's own incentives change too. The model can be wrong, and the last lesson looks at whether it is, but the mechanism is not a quirk of the numbers chosen here.

Example. In the original audit game, suppose better software cuts the cost of a wasted audit from 20 to 5. What happens to the two probabilities?

Only the authority's payoffs changed, so only the taxpayer's behaviour moves. The authority's indifference becomes 80p-5(1-p)=-50p, so 85p-5=-50p, giving 135p=5 and p=1/270.037. The taxpayer's indifference is unchanged, so the audit rate stays at q=1/3. Evasion falls from 13.3 per cent to 3.7 per cent, which is what the penalty increase failed to achieve.

Now you. Back to the original numbers, but the amount recovered by a successful audit rises from 80 to 130. What are the new equilibrium probabilities?

Answer

Again only the authority's payoffs move, so q stays at 1/3. The new indifference is 130p-20(1-p)=-50p, so 150p-20=-50p and 200p=20, giving p=0.1. Evasion falls from 13.3 per cent to 10 per cent. Anything that improves the return on enforcement reduces the offence; anything that only punishes the offender reduces enforcement.

Solving a two-by-two

The method is worth stating as a recipe, because most of the mixed equilibria anyone computes by hand are two-by-two.

Write down the opponent's payoffs. Set the opponent's expected payoff from their first action equal to their expected payoff from their second, as a function of your probability. Solve for your probability. Then swap the roles and repeat. Two linear equations, one unknown each, and no simultaneous solving is needed because the equations decouple.

Here is a version of matching pennies with unequal stakes. Player 1 wins if the coins match, and a match on heads is worth more than a match on tails. The game is zero-sum, so only player 1's payoffs are shown, and player 2's are their negatives:

Player 1headstails
Heads2-1
Tails-11

Let p be the probability player 1 plays Heads. Player 2's expected payoff from heads is -2p+1(1-p)=1-3p, and from tails it is p-(1-p)=2p-1. Equal when 5p=2, so p=0.4. Now let q be the probability player 2 plays heads. Player 1's expected payoff from Heads is 2q-(1-q)=3q-1, from Tails it is -q+(1-q)=1-2q, and these are equal when 5q=2, so q=0.4 as well. The value of the game to player 1 is 3(0.4)-1=0.2.

Look at what that says. The action with the bigger prize, Heads, is played less than half the time by the player who wants the match. Raising the reward for matching on heads makes player 1 play heads less often, because it is player 2 who reacts to that reward, by covering heads more, and player 1's own frequency is set by player 2's payoffs, which did not change until we changed them. The pattern is the same one the audit game showed, in a smaller game where it is harder to hide.

Example. Change the payoff for matching on heads from 2 to 3, leaving the rest. What is the new equilibrium, and what is the game worth to player 1?

Player 2's indifference: -3p+(1-p)=p-(1-p), so 1-4p=2p-1, giving p=1/3. Player 1's indifference: 3q-(1-q)=-q+(1-q), so 4q-1=1-2q and q=1/3. The value to player 1 is 4(1/3)-1=1/30.333. Player 1 now plays Heads only a third of the time and is better off than before, earning 0.333 rather than 0.2.

Now you. In the original version of this game (matching on heads worth 2), what is player 1's expected payoff if they stubbornly play Heads with probability 0.5 while player 2 plays the equilibrium q=0.4?

Answer

Player 2 is playing their equilibrium mixture, so player 1 is indifferent between Heads and Tails, each worth 0.2. Any mixture of them is therefore also worth 0.2, including 0.5. Deviating costs nothing, which is the indifference condition seen from the other side: an equilibrium mixture protects you against every deviation without punishing any of them.

Nash's theorem

Every finite game has at least one Nash equilibrium in mixed strategies. That is Nash's 1950 result, and it is what makes the concept usable, because a solution concept that frequently fails to apply is not a theory of anything.

The proof is a fixed-point argument and its shape is worth knowing even without the topology. Consider the map that takes any profile of mixed strategies and returns the set of best responses to it. A Nash equilibrium is exactly a profile that is in its own image, meaning a fixed point of that map. The set of mixed strategy profiles is a closed, bounded, convex set: a product of simplices. The best-response map is convex-valued, because indifference means any mixture of best responses is a best response, and it is well behaved in the technical sense the theorem requires. Kakutani's fixed point theorem then guarantees a fixed point exists. Nash's original note used exactly this; a year later he gave a shorter proof using Brouwer's theorem instead.

Two warnings come with it. The theorem promises existence, not uniqueness, and not that the equilibrium is easy to find: computing one is hard in a precise complexity-theoretic sense for large games. And it needs finiteness, or some substitute for it. Games with infinitely many strategies can fail to have equilibria without further assumptions, which matters two lessons from now when strategy sets become intervals of real numbers.

A related counting fact, due to Robert Wilson in 1971, is that almost every finite game has an odd number of equilibria. A two-by-two game with two pure equilibria therefore usually has a third one in mixed strategies, hiding between them, and finding that third one is a standard exercise in the next lesson.

What is a player actually doing?

Nobody believes a taxpayer tosses a fifteen-sided die. There are three defensible readings of a mixed equilibrium, and applied work should be explicit about which one is meant.

The first is literal randomisation, and it is exactly right in some places. Tennis players, poker players and penalty takers deliberately vary; tax authorities and customs officers really do sample at random; the routing of Allied convoys and the scheduling of security patrols are randomised by design. Where being predictable is fatal, mixing is a conscious operational choice.

The second is a population frequency. If a fresh taxpayer is drawn each year from a large population, p=0.133 can mean that 13.3 per cent of taxpayers evade with certainty and the rest never do, and the authority faces the same expected payoff. This is Nash's own mass action interpretation, and it is the natural reading in biology, where no animal randomises but a population can settle at a stable proportion of behaviours.

The third is Harsanyi's purification, from 1973, and it is the most satisfying. Suppose each player's payoffs are subject to small private fluctuations that only they observe: a taxpayer's mood, an auditor's caseload. Then almost every player has a strict pure best response given their own private shock, so nobody randomises, and yet the proportion choosing each action, as seen from outside, converges to the mixed equilibrium of the unperturbed game as the fluctuations shrink. Mixed equilibrium is then a description of an observer's ignorance rather than of a player's dice.

Where this leaves us

Existence is now settled, and the price of settling it is that uniqueness is gone for good. The audit game had exactly one equilibrium, but the three-by-three grid of the previous lesson had two, and any game whose players want to do the same thing as each other will have several.

That is not a technical nuisance. When a game has three equilibria, the theory as it stands says all three are consistent with rational play, and offers no reason to expect one rather than another. Whether anything can break the tie, and what the experiments say happens when real groups face exactly this problem, is the next lesson.