The chain store of the previous lesson unravelled because it had a twentieth town, and almost nothing that matters in trade, politics or family life has a last round anyone can name.
Why a known last round destroys everything
Take the prisoner's dilemma in its numerical form and play it exactly twice. Mutual cooperation pays 4 each, mutual defection pays 2 each, a sole defector takes 6 and leaves the cooperator 1, so in the notation of the dominance lesson , , and .
Solve it backwards, as the previous two lessons taught. In the second round there is no future to protect. Whatever happened in the first round, the players face a one-shot prisoner's dilemma in which defection strictly dominates, so both defect and each earns 2. That conclusion holds after every possible history, which is the fatal part: the second round pays the same whatever happened in the first. First-round play therefore cannot buy anything, and the first round is itself a one-shot prisoner's dilemma with defection strictly dominant.
The same argument runs for three rounds, for a hundred, for any finite number known to both players. The unique subgame perfect equilibrium of a finitely repeated prisoner's dilemma is defection in every round after every history. Repetition, by itself, achieves precisely nothing.
One caveat is worth stating because it is a genuine exception rather than a technicality. The unravelling needs the stage game to have exactly one Nash equilibrium. If the stage game has two, a good one and a bad one, then the final round can be used as a reward or a punishment, and cooperation becomes sustainable in the earlier rounds of a finite repetition. That is a real construction, and it is also a reminder that the result above is about the prisoner's dilemma in particular, not about finiteness in general.
Putting a number on the future
The escape is to remove the known end. Two quite different things make a payoff next period worth less than the same payoff now, and both are captured by one number.
The first is interest. A pound next year is worth today at interest rate . The second is mortality: the relationship may simply stop. If it continues to the next period with probability , then next period's payoff arrives only with that probability. Multiply them together and define the discount factor
so that a payoff received one period from now counts as today. Nothing in the analysis distinguishes the two sources, which is convenient: a relationship that might end and a player who is impatient are the same problem with the same arithmetic.
A constant flow of per period, starting now and continuing forever, is then worth
by the geometric series, which converges because . The same series has a second reading that is worth carrying around: with continuation probability and no interest, is the expected number of periods the relationship lasts. A discount factor of 0.9 means an expected ten more rounds; 0.99 means a hundred.
Example. A supplier earns £30,000 a year from a customer. The relationship survives each year with probability 0.9, and the supplier's cost of capital is 5 per cent. What is the relationship worth, and how long is it expected to last?
The discount factor is . The value is , which is seven years of revenue rather than the twenty a naive undiscounted count of a 10 per cent annual failure rate might suggest. The expected length is years, and the two numbers agree because at a zero interest rate the value would be exactly the flow times the expected length.
Now you. The same £30,000 flow, but the relationship survives each year with probability 0.8 and the cost of capital is 4 per cent. Find , the value and the expected length.
Answer
, so the value is and the expected length is years. A drop of ten points in the survival probability has cut the relationship's worth by nearly two fifths, which is the sense in which fragile relationships are cheap to betray.
Grim trigger and the patience threshold
Now build a strategy for the indefinitely repeated prisoner's dilemma. The simplest one that works is grim trigger: cooperate in the first round, and cooperate in every later round provided every previous round was mutual cooperation; if anyone ever defects, defect in every round thereafter, forever.
Suppose both players use it. On the path, both cooperate every round, so each earns
Now consider a player who defects in the first round instead. They collect that round, and from the next round onwards they face a permanent defector, so the best they can do is defect too and collect forever:
Cooperation is sustained when the first is at least the second. Multiply both sides by and the condition collapses to , which rearranges to . Dividing by , a negative number, flips the inequality:
Read the fraction. The numerator is what you gain by cheating for one round. The denominator is what you lose, per round, forever afterwards. Patience has to be large enough that the second outweighs the first, and the sucker payoff never appears, because a player following grim trigger is never caught out by a defection twice.
With the running numbers the threshold is . So a prisoner's dilemma played by players who expect at least one more round as likely as not becomes a game in which cooperation is an equilibrium. Notice what has and has not changed. The stage game still has defection strictly dominant; nobody has become nicer; there is no contract and no enforcement. What changed is that the strategy set now contains plans that condition on history, and one of those plans makes cooperation a best response to itself.
Grim trigger also passes the credibility test of the previous lesson. After a defection, both players defecting forever is the infinite repetition of the stage equilibrium, which is a Nash equilibrium of every subgame it starts. The threat is not merely announced, it is one the punisher is content to carry out.
Example. A stage game has , , and . What discount factor does grim trigger need, and what expected relationship length does that correspond to at a zero interest rate?
. At that means a continuation probability of at least one third, so an expected length of rounds. Cooperation here is cheap to sustain because the punishment is severe relative to the temptation: cheating gains 2 and costs 4 every round after.
Now you. Another stage game has , , and . Find the threshold and the expected length it implies.
Answer
, so the expected length must be at least rounds. The temptation is now larger than the whole cooperative payoff, so only a long relationship deters it. Raising alone always raises the threshold, which is why a one-off windfall available to a partner is the standard way relationships break.
The cartel's own threshold
The quantity-competition lesson left three numbers on the table, computed rather than assumed. In the Cournot market with demand and unit cost £40, the equilibrium gives each firm £400, the shared monopoly gives each £450, and a firm that cheats optimally on the cartel while its rival holds at 15 units earns £506.25. Those are , and .
The threshold follows directly:
Now do it in symbols, because the answer is more interesting than the arithmetic. Write for the gap between the demand intercept and marginal cost. The Cournot profit per firm is , the cartel share is , and the optimal cheat against a rival producing is units at a margin of , worth . Substituting,
Every parameter has cancelled. In any symmetric linear-demand Cournot duopoly with constant marginal cost, the patience needed to hold the full cartel by grim trigger is exactly , whatever the size of the market, the slope of demand or the cost. That is a genuinely checkable prediction and an uncomfortable one: at a zero termination risk it corresponds to an interest rate of up to per cent per period, so any duopoly reviewing prices monthly should collude effortlessly.
Real cartels are not effortless. They form, cheat, break and reform, which means the model is missing something, and the missing piece is usually the assumption that a defection is seen immediately. Suppose cheating takes periods to detect, so the cheat collects for periods before punishment starts. The same algebra gives
and since , raising raises the required patience sharply. For the cartel above, a two-period blind spot needs and a three-period one needs . The first of those is worth checking in full: at the cooperative stream is , and cheating for two periods before punishment yields , exactly indifferent.
The moral is that what protects a cartel is not the ferocity of the punishment but the speed of detection, which is why price transparency, most-favoured-customer clauses and published tariffs are treated by competition authorities as collusive devices rather than as conveniences for the customer.
Tit for tat, and the price of forgiveness
Grim trigger works, and nobody would want to use it. One mistaken keystroke, one misread invoice, and a profitable relationship is over forever. The obvious repair is to punish briefly and then forgive, which is tit for tat: cooperate in the first round, then copy whatever the opponent did last round.
Tit for tat needs more patience than grim trigger, and the reason is exactly its forgiveness. Two deviations have to be deterred rather than one.
Against a tit-for-tat opponent, a player who defects forever collects once and from then on, which gives the earlier condition . But there is a second and cheaper deviation: alternate. Defect, then cooperate while the opponent retaliates, then defect again as they forgive. That yields , , , , and so on, worth . Requiring the cooperative stream to beat it gives , that is . Tit for tat therefore sustains cooperation only when
With , , , the two thresholds are 0.5 and , so tit for tat needs where grim trigger needed 0.5. Forgiveness is not free: a punishment that ends is a punishment that can be walked into deliberately.
There is a further honesty about tit for tat that is often skipped. It is not subgame perfect. After a single defection, the two players enter an alternating pattern of defection and cooperation which is not an equilibrium of the subgame it starts, so the strategy relies on carrying out a punishment that the punisher would prefer to abandon. Its reputation comes from tournament performance rather than from the equilibrium arithmetic, and that distinction is the subject of the next lesson.
Example. With , , , and , does grim trigger sustain cooperation? Does tit for tat?
Cooperating forever is worth . Against grim trigger, defecting yields , which is less than 10, so cooperation holds. Against tit for tat, the alternating deviation yields , which beats 10, so it does not. At a pair of grim triggers cooperate and a pair of tit-for-tatters are exploited every other round.
Now you. Take , , , and . Check both strategies.
Answer
Cooperation is worth . Permanent defection yields , and the alternating deviation yields . Both fall short of 24, so both strategies sustain cooperation here. In threshold form, grim needs and tit for tat needs the larger of that and , and 0.75 clears both.
What repetition actually needs
The result is strong enough to be misapplied, so it is worth naming the conditions that carry it.
Defections must be observed. Everything above assumes each player sees what the other did. When the signal is noisy, a low price may mean a cheating rival or a weak market, and punishing on the evidence means punishing the innocent. Edward Green and Robert Porter showed in 1984 that the best arrangement under noisy monitoring has price wars occurring on the equilibrium path, triggered by bad demand rather than by anyone's misconduct. Porter's 1983 study of the Joint Executive Committee, the American railroad cartel of the 1880s, found exactly that pattern of periodic collapses in a cartel that was otherwise holding.
There must be no known end. A fixed and commonly known horizon restores the unravelling of the first section. What matters is not that the relationship is literally infinite but that at every point there is a future worth protecting.
The same parties must meet again, and be recognised. Repetition works through identity. Anonymity dissolves it, which is why online markets invest so heavily in making a seller's history follow them and why one-shot tourist trades in every country are the ones with the worst prices.
Nobody has become nicer. Repetition changes the payoffs of the strategies, not the preferences of the players. That also means it is morally neutral: the same argument that sustains a technical standard sustains a price-fixing ring, an omerta, and a cycle of retaliation between neighbours. The theory explains why cooperation is possible, not why it is good.
Where this leaves us
The finite prisoner's dilemma had exactly one equilibrium outcome and it was bad. Adding a shadow of the future has produced a second one, mutual cooperation, above a patience threshold that can be computed from a cartel's own accounts.
It has produced rather more than that, and the excess is a problem. Nothing in the argument was special to mutual cooperation: the same construction, with a suitable punishment behind it, supports alternating exploitation, a ninety-ten split of the gains, or a deliberately wasteful arrangement that neither player likes. The next lesson states exactly how much the repeated game can support, and the answer is close to everything, which turns a triumph into a crisis about what the theory is still predicting.