Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Mathematical Foundations

Rebuild school mathematics as an adult: algebra as structure, functions and graphs, logarithms and exponentials, trigonometry, and a first taste of proof.

From counting to the real line

Mathematics is usually taught as though the numbers were simply there, waiting, and the only question was what to do with them. That is backwards: every kind of number was invented under protest, because someone wrote down a problem that the numbers already available could not answer.

This lesson follows that chain of protests from one end to the other. It assumes nothing beyond arithmetic, and what it produces is the object the rest of the course lives on, the real line, along with an honest account of what that line still cannot do.

Counting, and the first thing it cannot do

Start with what a child starts with: 1,2,3,4,dots, the natural numbers, written N. They answer one question, "how many", and they are closed under two operations. Add two naturals and you get a natural. Multiply two naturals and you get a natural. Closure is worth naming, because it is exactly what fails next.

Now ask a question of the same shape as one they can answer. The equation x+3=7 has the answer 4, and nothing is wrong. The equation x+7=3 has no answer at all in N. It is not that the answer is hard to find; there is no natural number that works, so within this system the question is malformed.

There are two honest responses. One is to declare the question illegal and keep the numbers small and safe, which is roughly what Greek mathematics did. The other is to invent whatever object the equation demands and then check that the new object does not wreck the arithmetic that already worked. Every step in this lesson is that second response, taken again and again.

Zero and the negatives

The demand of x+7=3 is for a number that undoes addition. Grant it: for every natural n introduce an object -n with the defining property n+(-n)=0, and introduce 0 itself as the number that changes nothing when added. The result is the integers, Z, running dots,-2,-1,0,1,2,dots and closed under addition, subtraction and multiplication.

The first systematic account of how these behave is Brahmagupta's, in the Brahmasphutasiddhanta of 628 CE, which states rules for zero and for what he calls fortunes and debts: a debt subtracted from zero is a fortune, and the product of two debts is a fortune. That second rule is the one every student finds arbitrary, and it is not arbitrary at all. It is forced, and the next lesson shows exactly what forces it.

Resistance lasted a very long time. Cardano, solving cubics in the Ars Magna of 1545, still called negative roots fictae, fictitious, and European mathematicians were arguing about whether negative quantities were legitimate into the eighteenth century. Two facts sit behind this. A negative number answers no question of the form "how many", so it has no direct counterpart in a heap of stones. And a debt, a temperature below freezing, or a displacement to the left are all perfectly concrete, so the objection was never about the world. It was about what a number was allowed to be.

Fractions, and a system that looks complete

Multiplication now has the same problem addition had. The equation 3x=12 is fine. The equation 3x=7 has no solution among the integers, and the fix is the same trick: for every non-zero integer n introduce an object 1/n with the defining property n×(1/n)=1. What comes out is the rational numbers, Q, every number expressible as p/q with p and q integers and q not zero.

The rationals are closed under all four operations, division by zero excepted, and they are also dense: between any two distinct rationals there is another, since their average is rational and sits strictly between them. Take 1/3 and 2/5; their average is 11/30, which lies between the two, and the same move repeats forever. There is no such thing as the next rational after a given one.

Density makes the system feel finished. Pick any point on a ruler and you can name a rational as close to it as you please, so it is natural to assume the rationals fill the line completely. That assumption is false, and the discovery that it is false is the most important event in this lesson.

Every rational has a decimal expansion that either terminates or eventually repeats, and every terminating or repeating decimal is a rational. The forward direction is long division: dividing by q can only produce q possible remainders, so a remainder must recur, and once it does the digits cycle. The reverse direction is a trick worth having.

Example. Write 0.272727dots as a fraction.

Call the number x. The repeating block is two digits long, so multiply by 100 to shift the expansion by exactly one block: 100x=27.272727dots. Subtracting the original, the infinite tails cancel exactly, since they are identical: 100x-x=27, so 99x=27 and x=27/99. Dividing top and bottom by 9 gives x=3/11. Checking by division, 3 divided by 11 is 0.2727dots, as claimed.

Now you. Write 0.135135dots as a fraction in lowest terms.

Answer

The block is three digits, so 1000x-x=135, giving x=135/999. Both are divisible by 27, since 135=27×5 and 999=27×37, so x=5/37.

The diagonal that is not a fraction

Draw a square of side 1 and ask for the length d of its diagonal. Pythagoras' theorem gives d2=12+12=2, so d is a number whose square is 2. The Pythagoreans, working in the fifth century BCE, could prove that no fraction has this property. Aristotle refers to the argument as a familiar one, and a version of it survives in Book X of Euclid's Elements.

The full proof is the business of the last lesson of this course, where the technique it uses is the point. What matters here is the consequence, and it is severe. The diagonal of a perfectly ordinary square has a length, you can draw it, and that length is not any ratio of whole numbers. The rationals, dense as they are, have a hole in them exactly where the diagonal lands.

The tradition that the discovery was a scandal, and that Hippasus was drowned at sea for divulging it, is a late story and probably not history. The mathematical damage, however, was real. A school whose slogan was that all is number had assumed that any two lengths share a common unit small enough to measure both, and the diagonal of a square shows they need not.

You can watch the failure numerically. Fractions approximate 2 beautifully, and never hit it.

Example. How close does 99/70 come to having a square of 2?

Square it: 992=9801 and 702=4900, so (99/70)2=9801/4900. Now 2×4900=9800, so the square is 9801/4900=2+1/4900, missing 2 by exactly one part in 4900, about 0.0002. Close, and not equal, and the argument of the last lesson says the miss can never be zero however clever the fraction.

Now you. By how much does the square of 577/408 miss 2?

Answer

5772=332929 and 4082=166464, and 2×166464=332928, so the square is 332929/166464=2+1/166464. It misses by exactly one part in 166464, which is roughly six millionths, and still misses.

The real line

Patch the holes and you get the real numbers, R: every point on a continuous line has a number, and every number a point. Informally, a real number is a decimal expansion, allowed to run forever without repeating. The rationals are the expansions that terminate or repeat; the irrationals are all the rest, and 2=1.41421356dots is one of them.

Making that respectable took until the 1870s, when Dedekind and Cantor gave constructions of the reals in terms of the rationals alone, so that the line stopped being a picture and became a definition. The property their constructions deliver is completeness: there are no gaps left, and any quantity you can squeeze arbitrarily tightly between rationals is itself a real number. Completeness is what makes limits work, so it is the foundation the whole of calculus is built on, and it is why the course you are reading stops here and calculus starts here.

Two facts about the irrationals are worth carrying. First, they come in two kinds. 2 is algebraic, meaning it solves a polynomial equation with whole-number coefficients, here x2-2=0. Numbers that solve no such equation at all are transcendental, and both of the famous constants are: Hermite proved it for e in 1873 and Lindemann for π in 1882. Lindemann's result settled squaring the circle, an open problem of two thousand years, in the negative.

Second, the irrationals are not rare. Cantor showed in 1874 that the rationals can be listed in an infinite sequence while the reals cannot, so the two infinities differ in size, and in a precise sense almost every real number is irrational. The numbers we can name are the exceptions.

Decimals, precision and error

In practice, every real number that is not a whole number gets replaced by a decimal that stops, so it is worth being able to say how much that costs. The absolute error is the difference between the approximation and the true value; the relative error is that difference divided by the true value, and it is the one that usually matters, because being wrong by a metre is trivial for a road and fatal for a doorframe.

Example. π=3.14159265dots and the schoolroom approximation is 22/7=3.14285714dots. What are the absolute and relative errors?

Subtracting, the absolute error is 3.14285714-3.14159265=0.00126449, so 22/7 is a little too large. Dividing by π gives a relative error of 0.00126449/3.14159265=0.000402, about four parts in ten thousand, or 0.04%. On a circle the size of a dinner plate, say 26 cm across, the circumference is about 817 mm, so that is an error of roughly a third of a millimetre.

Now you. The better approximation 355/113=3.14159292dots was known to Zu Chongzhi in the fifth century. What is its absolute error, and roughly its relative error?

Answer

The absolute error is 3.14159292-3.14159265=0.00000027, near enough 2.7×10-7. Dividing by π gives a relative error of about 8.5×10-8, under one part in ten million, which is why the approximation stood as the best known for nearly a thousand years.

The equation the line cannot solve

The chain of this lesson has a shape: name an equation with no solution, invent the solution, check that nothing already working breaks. Run it once more and it does not terminate.

The equation is x2=-1. On the real line it has no solution, and for a reason stronger than mere absence: any real number, positive or negative, has a square that is zero or positive, so no arrangement of real numbers will ever produce -1. The same demand can be made anyway, and it was, by Bombelli in 1572, who found that solving cubic equations sometimes required carrying square roots of negatives through the middle of a calculation even when the final answers were ordinary whole numbers. Grant the demand and you get the complex numbers, and everything that worked before still works.

This course does not follow that road, and the reason is worth saying plainly rather than hiding. The material ahead, functions, graphs, growth, decay and the geometry of angles, is about quantities you can measure and plot: a length, a temperature, a balance, a height. Those live on the line. The complex numbers are the natural home of a different set of questions, and taking them on here would double the machinery without touching the goal.

So the working set is R, with N, Z and Q sitting inside it, and one honest gap left at x2=-1.

There is something conspicuously unfinished about all of this. Solving x+7=3 and 3x=7 meant subtracting from both sides and dividing both sides, and no reason was ever given for why those moves are allowed. They were used because they are familiar. The rules that licence them are few, they can be written down completely, and once they are, results such as Brahmagupta's rule about two debts stop being conventions to memorise and become consequences. That is the next lesson.

The rules of algebra

Solving x+7=3 meant subtracting seven from both sides, and nobody said why that was allowed. The move is so familiar that the question sounds pedantic, but the whole of algebra is built out of a handful of such moves, and if they are taken on faith then every result downstream rests on faith too.

This lesson writes the licence down. It assumes only the number systems of the previous lesson, the naturals inside the integers inside the rationals inside the reals, and it produces a list short enough to memorise from which everything else follows, including the rule about two negatives that every student is told to accept.

The moves nobody justified

School algebra is taught as a set of permissions handed out one at a time. You may add the same thing to both sides. You may multiply out a bracket. You may cancel a common factor. You may not divide by zero. Each arrives as a separate instruction, and the impression left is of a large body of rules held together by nothing.

The truth is the reverse. There are about eight facts about addition and multiplication, all of them things you would guess, and every legitimate manipulation is a consequence of them. Nothing else is needed and nothing else is allowed. When a manipulation goes wrong, and one will go wrong later in this lesson in a way that produces the proof that 1=2, the fault is always that some step used a law outside the conditions the law actually carries.

The name for a system carrying these laws is a field, and Q and R are both fields. The integers are not, because they lack multiplicative inverses: 3x=7 has no integer solution, which is exactly the failure that forced the rationals into existence.

The list

Addition is commutative, a+b=b+a, and associative, (a+b)+c=a+(b+c). It has an identity, the number 0 with a+0=a, and every a has an additive inverse -a with a+(-a)=0. Multiplication has the same four properties: ab=ba, (ab)c=a(bc), the identity 1 with a×1=a, and for every a except zero a multiplicative inverse a-1 with a×a-1=1.

That is four laws twice over, and they say nothing about how the two operations interact. One law does: distributivity,

a(b+c)=ab+ac

This is the only bridge between addition and multiplication in the whole list, and it is the busiest law in mathematics. Expanding brackets is distributivity read left to right; factoring is the same law read right to left; long multiplication of ordinary numbers is distributivity applied to place value.

The laws about equations follow from these plus one fact about equality: if a=b, then a and b are the same number, so anything true of one is true of the other. Adding c to both sides is legal because a+c and b+c are the same number when a and b are. That is the licence that was missing in the previous lesson, and it costs one sentence.

Subtraction and division are shorthand

Notice what the list does not contain. There is no law of subtraction and no law of division, and this is not an oversight. Subtraction is defined out of what is there: a-b means a+(-b), the addition of the additive inverse. Division likewise means multiplication by the multiplicative inverse, a÷b=a×b-1.

This is worth more than tidiness, because it explains a fact every student meets as a nuisance. Addition is commutative and associative; subtraction is neither. 7-3 is not 3-7, and (10-4)-3=3 while 10-(4-3)=9. The rules did not fail. Subtraction was never on the list, and the operation it abbreviates is commutative only in the sense that a+(-b)=(-b)+a, which is not the statement that swapping the visible numbers is safe.

The same holds for division, which is why 12÷4÷2 is ambiguous without a convention, while 12×4-1×2-1 is not ambiguous at all. Reading a subtraction as the addition of a negative, and carrying the sign with the number rather than leaving it stranded between two terms, removes most sign errors permanently.

What the laws force

The rule that a negative times a negative gives a positive is presented to most people as a convention, and Brahmagupta's fortunes and debts do read like a convention. They are not. Given the laws above, the rule is the only possibility, and here is the argument.

First, anything times zero is zero. Since 0+0=0, multiply both sides by a and distribute: a×0=a(0+0)=a×0+a×0. Now add -(a×0) to both sides, which is legal, and the left becomes 0 while the right becomes a×0. So a×0=0, derived rather than assumed.

Now take (-1)×(-1). Start from (-1)+1=0 and multiply through by -1:

(-1)big[(-1)+1big]=(-1)×0=0

Distribute the left side: (-1)(-1)+(-1)(1)=0, which is (-1)(-1)+(-1)=0. Add 1 to both sides and (-1)(-1)=1. If a negative times a negative were anything else, distributivity would fail, and with it every bracket anyone has ever expanded. The rule is not a choice made for convenience; it is the price of keeping the arithmetic consistent.

Example. Use distributivity to compute 97×103 mentally.

Write it as (100-3)(100+3) and expand: 100×100+100×3-3×100-3×3. The two middle terms cancel exactly, leaving 10000-9=9991. This is the difference of two squares, (a-b)(a+b)=a2-b2, and it is nothing but distributivity applied twice.

Now you. Compute 58×62 the same way.

Answer

Both numbers sit two from 60, so this is (60-2)(60+2)=3600-4=3596.

Why zero has no inverse

The one prohibition in school algebra is the ban on dividing by zero, and it is usually taught as a rule with a vague warning attached. It is a theorem. The laws say every number except zero has a multiplicative inverse, and the exception is forced, because granting zero an inverse contradicts a result already proved.

Suppose 0-1 existed, meaning some number c with 0×c=1. But it was proved two sections ago that a×0=0 for every a, so 0×c=0, and therefore 0=1. Once zero and one are the same number, every number is zero, since a=a×1=a×0=0, and arithmetic collapses into a single point. Division by zero is not undefined out of squeamishness. Defining it destroys the system.

This is also where the famous fake proof that 1=2 hides. Start with a=b, multiply by a to get a2=ab, subtract b2 to get a2-b2=ab-b2, factor both sides into (a-b)(a+b)=b(a-b), then cancel the common factor (a-b) to get a+b=b. With a=b=1 that reads 2=1. Every step is legal except the cancellation, which is multiplication by (a-b)-1, and a-b is zero. Cancelling a common factor is only ever multiplication by an inverse, so it always carries the condition that the factor is not zero, and forgetting that condition is the commonest way to lose a solution to an equation.

Distributivity in both directions

Expanding is mechanical: every term in the first bracket multiplies every term in the second. Factoring is the same law used backwards, and it is harder because it is a search rather than a procedure. It is worth the effort because a factored expression displays its zeros: a product is zero exactly when one of its factors is, which is the fact the whole of equation solving rests on.

For a quadratic x2+bx+c the search has a shape. If it factors as (x+p)(x+q), then expanding gives x2+(p+q)x+pq, so p and q are two numbers that multiply to c and add to b. Since there are finitely many integer pairs multiplying to c, the search terminates.

Example. Expand (2x+3)(x-5), then factor x2-x-12.

Expanding: 2x×x+2x×(-5)+3×x+3×(-5)=2x2-10x+3x-15=2x2-7x-15. Factoring: two numbers multiplying to -12 and adding to -1 are -4 and 3, so x2-x-12=(x-4)(x+3). Check by expanding: x2+3x-4x-12=x2-x-12.

Now you. Expand (3x-4)(2x+5), then factor the result back.

Answer

6x2+15x-8x-20=6x2+7x-20. Factoring back means finding the same pair of brackets, (3x-4)(2x+5), which is the only factorisation with integer coefficients.

Order of operations, and what is actually a convention

Some of what looks like a law is only notation. The agreement that multiplication binds tighter than addition, so 2+3×4 is 14 and not 20, is a convention about reading, not a fact about numbers. It could have gone the other way, and expressions would be written with more brackets. Associativity, by contrast, is a fact: (2+3)+4 and 2+(3+4) are equal whatever anyone agrees.

Distinguishing the two matters when a manipulation feels forbidden. Rearranging a+bc into bc+a is licensed by commutativity of addition. Rewriting it as (a+b)c is licensed by nothing and is false, as 2+3×4=14 against (2+3)×4=20 shows. When in doubt, put in the brackets the convention allowed you to omit, and check which law the intended step is appealing to.

Example. Simplify 5(2x-3)-2(4x-7).

Distribute each bracket, taking the second sign with the 2: 10x-15-8x+14. The -2×-7=+14 is the negative-times-negative rule proved above. Collecting like terms, which is distributivity backwards on 10x-8x=(10-8)x, gives 2x-1.

Now you. Simplify 3(4x+5)-4(2x-3).

Answer

12x+15-8x+12=4x+27.

Where the laws stop

Two honest limits are worth stating. First, the laws describe addition and multiplication only. Inequalities carry their own rules, and one of them famously breaks the pattern: multiplying both sides of a<b by a negative number reverses it, since 2<3 but -2>-3. Nothing in the field laws covers order, and an ordered field is a stronger structure than a field.

Second, commutativity is not a property of everything worth multiplying. It fails for rotations in space, for the matrices met in linear algebra, and for the operators of quantum mechanics, where the failure is the physics. That such systems exist is the reason the laws are stated explicitly rather than assumed: writing them down is what makes it possible to notice which ones a new system keeps and which it abandons.

Within R, though, the list holds without exception, and it licences every rearrangement the rest of this course performs. Rearranging is only half of what algebra does. The other half begins by taking two expressions and asserting they are equal, which turns a manipulation into a question with an answer, and answering that question is the next lesson.

Equations and the quadratic

An expression such as 2x2-7x-15 is not true or false; it is just a recipe waiting for a number. Set it equal to something and it becomes a claim about x, and the claim is true for some values and false for others. Finding exactly which ones is what solving means.

This lesson takes that from the easiest case to the first genuinely hard one. It assumes the laws of the previous lesson, since every step in a solution is one of them applied to both sides, and it ends with a formula for every quadratic that is derived here rather than quoted.

From expression to statement

An equation asserts that two expressions name the same number. Its solutions are the values of the unknown that make the assertion true, and to solve it is to replace it with a chain of equations having the same solutions until the last one reads x=something.

The chain works because of the equality fact from the previous lesson: doing the same thing to two names for one number leaves two names for one number. Adding, subtracting, and multiplying by anything non-zero all preserve the solution set exactly, so nothing is gained or lost.

Two operations do not preserve it, and both are traps worth naming now. Multiplying by an expression that might be zero can invent solutions, since x=3 has one solution but x(x-5)=3(x-5) has two. Squaring both sides does the same, because x+6=x becomes x+6=x2, whose solutions are 3 and -2, and -2 fails in the original since the square root sign means the non-negative root. The fix is not to avoid these moves but to check every candidate in the original equation, which costs seconds.

Linear equations by inverse operations

A linear equation has the unknown to the first power only, and it always yields to the same strategy: undo the operations applied to x, in reverse order. Collect all the x terms on one side, all the numbers on the other, and divide by whatever multiplies x.

Example. Solve 5(x-3)+2=3x+7.

Distribute: 5x-15+2=3x+7, so 5x-13=3x+7. Subtract 3x from both sides: 2x-13=7. Add 13: 2x=20. Divide by 2: x=10. Check in the original, which is not optional: the left is 5×7+2=37 and the right is 30+7=37.

Now you. Solve 4(2x+1)-3=2(x-4)+11 and check it.

Answer

8x+1=2x+3, so 6x=2 and x=1/3. Checking, the left is 4(2/3+1)-3=20/3-3=11/3 and the right is 2(1/3-4)+11=-22/3+11=11/3.

Not every linear equation has exactly one solution. Reduce 2(x+3)=2x+6 and the unknown vanishes, leaving 6=6, true for every x: the equation was an identity in disguise. Reduce 2(x+3)=2x+5 and you get 6=5, false for every x, so there is no solution. A vanishing unknown is information, not a mistake, and which of the two you have is decided by whether what remains is true.

Quadratics, and why factoring is not enough

A quadratic equation is ax2+bx+c=0 with a not zero. When the left side factors over the integers the work is already done, because a product is zero exactly when one factor is: from (x-4)(x+3)=0 read off x=4 and x=-3 immediately.

The method is fast and it is unreliable, because most quadratics do not factor over the integers. The equation x2-6x+2.5=0 has perfectly good solutions and no integer pair multiplies to 2.5 while adding to -6. Searching for a factorisation that does not exist can take a long time before you conclude anything, so a method that always works is worth having.

The one that always works comes from noticing which quadratics are easy. Anything of the form (x+p)2=k is solved in one line: take square roots, remembering both signs, so x+p=±k and x=-p±k. So the task is to turn any quadratic into that shape, and the technique for doing so is called completing the square.

Completing the square

The identity to exploit is (x+p)2=x2+2px+p2. Given x2+bx, matching the middle terms gives 2p=b, so p=b/2, and the perfect square that starts with those two terms is (x+b/2)2=x2+bx+b2/4. It has an extra b2/4 on the end, so subtract it back off: x2+bx=(x+b/2)2-b2/4. That single line is the whole technique.

Example. Solve 2x2-12x+5=0 exactly, then to four decimal places.

Divide through by 2 so the leading coefficient is 1: x2-6x+2.5=0. Half of -6 is -3, so x2-6x=(x-3)2-9, and the equation becomes (x-3)2-9+2.5=0, that is (x-3)2=6.5. Take roots: x=3±6.5. Since 6.5=2.5495, the solutions are x=5.5495 and x=0.4505. Check the first: 2(5.5495)2-12(5.5495)+5=61.594-66.594+5=0 to the precision shown.

Now you. Solve 3x2+12x-7=0 by completing the square, exactly and to four decimals.

Answer

Divide by 3: x2+4x-7/3=0. Half of 4 is 2, so (x+2)2-4-7/3=0 and (x+2)2=19/3. Then x=-2±19/3, and 19/3=2.5166, giving x=0.5166 and x=-4.5166.

The formula, derived

Run that procedure once on the general equation and you never have to run it again. Start from ax2+bx+c=0 and divide by a, which is legal because a is not zero:

x2+bax+ca=0

Complete the square on the first two terms, with half the middle coefficient being b/2a:

(x+b2a)2-b24a2+ca=0

Move the constants across and put them over the common denominator 4a2:

(x+b2a)2=b2-4ac4a2

Take square roots of both sides. The right-hand denominator is a perfect square, so its root is 2a, and

x=-b±b2-4ac2a

That is the quadratic formula, and it is not a fact to be memorised on authority: it is the completed square, done once with letters instead of numbers. Anyone who forgets it can rederive it in five lines. Check it against the worked example above, where a=2, b=-12, c=5: the formula gives x=(12±144-40)/4=(12±104)/4, and 104=10.198, so x=5.5495 or 0.4505 as before.

The discriminant tells you before you start

The quantity under the root, b2-4ac, is the discriminant, and it decides the character of the answer before any arithmetic. If it is positive there are two distinct real solutions. If it is zero the ± collapses and there is one, a repeated root at x=-b/2a. If it is negative there are no real solutions at all, because no real number squares to a negative, which is precisely the gap left open in the first lesson at x2=-1.

So x2+4x+7=0 can be dismissed in one step: 16-28=-12, negative, no real solutions. This is not a failure of technique. The equation asks for a number whose square plus four times itself is -7, and on the real line there is none. The complex numbers supply two, and this course leaves them alone.

A repeated root is the boundary case between the two, and it is what a projectile problem gives when the object just grazes the height asked about. Discriminant zero is the algebraic signature of tangency, a fact that becomes visible as a picture in the next lesson but one.

A quadratic that came from somewhere

Formulas earn their keep on real problems. Throw a stone upward at 15 m s⁻¹ from a cliff 40 m above the sea. Ignoring air resistance, its height in metres after t seconds is h=40+15t-4.9t2, where 4.9 is half the gravitational acceleration of 9.8 m s⁻².

Example. When does the stone hit the water?

Set h=0: -4.9t2+15t+40=0, or 4.9t2-15t-40=0 after multiplying by -1. The discriminant is 225+4×4.9×40=225+784=1009, whose root is 31.765. So t=(15±31.765)/9.8, giving t=4.772 s and t=-1.711 s. Both are solutions of the equation; only one is a solution of the problem, since the stone was not in flight before it was thrown. Discarding a mathematically valid root on physical grounds is a normal part of modelling, and it should be done explicitly rather than silently.

Now you. A ball is thrown upward at 20 m s⁻¹ from a 25 m balcony, so h=25+20t-4.9t2. When does it land?

Answer

4.9t2-20t-25=0, discriminant 400+490=890, root 29.833. So t=(20+29.833)/9.8=5.085 s, the negative root being rejected.

The same quadratic answers other questions about the flight. The stone is highest when the two roots of h=H coincide, which is when the discriminant vanishes, and by symmetry that is halfway between the roots of any horizontal cut: t=15/9.8=1.531 s, at which h=51.48 m.

What the formula reveals, and cannot do

Look at the formula rather than through it. The solutions are expressed in terms of a, b and c, so as those coefficients vary the answers vary with them, and this is the first time in the course that a quantity has been described as depending on other quantities in a way no single equation captures. That dependence is the subject of the next lesson.

Two honest limits close this one. The formula is exact but not always the best way to compute: when b2 is enormously larger than 4ac, subtracting two nearly equal numbers in the numerator loses precision, and numerical libraries compute one root by the formula and the other from the fact that the roots multiply to c/a. And it does not generalise as far as one would hope. There is a formula for the cubic, published by Cardano in 1545, and one for the quartic, but Abel proved in 1824 that no formula in radicals exists for the general fifth-degree equation. Solving by formula stops at degree four, permanently.

Functions and inverses

The quadratic formula says that the solutions of ax2+bx+c=0 depend on the three coefficients, and that dependence is not itself an equation. It is a rule: hand it three numbers and it hands back an answer, every time, without ambiguity.

Mathematics has a name for such a rule and a machinery for handling it, and building that machinery is what this lesson does. It assumes only the algebra of the previous two lessons, and it produces the object that the rest of the course studies almost exclusively.

What a function is, exactly

A function is a rule that assigns to each input exactly one output. Three things must be given before the rule is fully specified: the set of allowed inputs, called the domain; the set the outputs are drawn from, the codomain; and the rule itself. Writing f(x)=3x+1 gives only the third, and the other two are usually left implicit as "whatever real numbers make sense".

The demand for exactly one output is the whole of the definition and it is stricter than it looks. The rule "y is a number whose square is x" is not a function, because x=9 yields both 3 and -3. This is why 9 is defined to mean 3 and not -3: the square root symbol has to name one number, and the choice of the non-negative one is a convention adopted to rescue functionhood. It is also why the quadratic formula carries an explicit ± rather than hiding two answers inside one symbol.

The notation f(x) is worth reading carefully, because it is not multiplication. It means the output of f at the input x, and f alone is the name of the function while f(x) is the name of a number. Sloppiness here is harmless in easy cases and disastrous when functions start being fed to other functions, which happens later in this lesson.

Domain: where the rule refuses to run

Two things stop a formula in its tracks over the reals, both met already. Division by zero is barred, so f(x)=1/(x-2) has domain every real except 2. Even roots of negatives do not exist on the real line, so g(x)=x-4 needs x-40, giving domain x4. Everything else is allowed, and a polynomial has every real number in its domain.

Domain can also be imposed from outside rather than discovered from the formula. If A(r)=πr2 is the area of a circle of radius r, then A(-3)=9π is arithmetically fine and physically meaningless, so the domain is r>0 because of what the function is for. Modelling almost always narrows a domain this way, and forgetting to say so is how a model gets used outside the range where it was ever true.

The range is what actually comes out: the set of values f(x) takes as x runs over the domain. Range is generally harder to find than domain, because it requires knowing the behaviour of the whole rule rather than checking for prohibited operations. For f(x)=x2 with domain all reals, the range is y0, since squares are never negative and every non-negative number is a square.

Example. Give the domain and range of g(x)=x-4+1.

The root demands x4, so that is the domain. As x runs from 4 upward, x-4 runs from 0 upward, its square root runs from 0 upward taking every non-negative value, and adding 1 shifts all of it, so the range is y1. A check: g(13)=9+1=4, and g(20)=16+1=5.

Now you. Give the domain and range of h(x)=3-x+5.

Answer

The domain is x-5. The root takes every value from 0 upward, and subtracting it from 3 takes 3 downward without limit, so the range is y3. Check: h(-5)=3 and h(4)=3-3=0.

A function need not be a formula

The formula is the commonest presentation but not the definition, and holding the two apart pays off. A table of values is a function if no input appears twice with different outputs. The mapping from a date to the closing price of an index is a function with no formula at all. A rule given in words, such as "round to the nearest integer", is a function.

Piecewise definitions matter in practice because so much of the world is written in bands. United Kingdom income tax in the 2024/25 year charges nothing on the first 12{,}570 pounds, 20 per cent on income between there and 50{,}270, and 40 per cent from there to 125{,}140. That is one function of income, defined by different formulas on different stretches.

On an income of 30{,}000 pounds the tax is 0.20×(30{,}000-12{,}570)=3{,}486 pounds. On 60{,}000 it is 0.20×(50{,}270-12{,}570)+0.40×(60{,}000-50{,}270)=7{,}540+3{,}892=11{,}432 pounds. Note what the calculation shows: the higher rate applies only to the part of the income above the threshold, so crossing a band boundary never reduces take-home pay. That single fact, which is a property of the function, is the one most often got wrong in public argument about tax.

Composition, and why order matters

Feeding the output of one function into another gives a new function, written (fg)(x)=f(g(x)) and read "f after g". The inner function runs first, which is the reverse of the reading order, and it is a common source of error.

Take f(x)=3x+1 and g(x)=x2. Then f(g(2))=f(4)=13, while g(f(2))=g(7)=49. Composition is not commutative, and unlike the commutativity of multiplication this failure is normal rather than exotic: doubling then adding three is a different rule from adding three then doubling, as anyone who has applied a discount before or after tax knows.

Composition also inherits domain restrictions from both parts. In f(g(x)) the input must be in the domain of g, and the value g(x) must then be in the domain of f. With g(x)=x-4 and f(u)=u, the composition needs x4 even though each function separately is happy with any real. Checking both conditions, not just the visible one, is the discipline.

Example. With f(x)=2x-5 and g(x)=x2+1, find f(g(3)), g(f(3)), and a formula for f(g(x)).

g(3)=10, so f(g(3))=20-5=15. In the other order, f(3)=1, so g(f(3))=1+1=2. In general f(g(x))=2(x2+1)-5=2x2-3, and a check at x=3 gives 18-3=15, agreeing.

Now you. With the same two functions, find a formula for g(f(x)) and evaluate it at x=4.

Answer

g(f(x))=(2x-5)2+1=4x2-20x+26. At x=4 that is 64-80+26=10, which matches f(4)=3 then g(3)=10.

Running a function backwards

An inverse of f is a function f-1 that undoes it: f-1(f(x))=x for every x in the domain, and f(f-1(y))=y for every y in the range. The superscript is unfortunate notation, since f-1 does not mean 1/f, and only context distinguishes them.

Not every function has one, and the reason is the same strictness that defined functions in the first place. If two different inputs give the same output, the inverse would have to send that output back to both, and it would not be a function. So f is invertible exactly when it is injective, meaning distinct inputs always give distinct outputs. The function f(x)=x2 on all the reals is not injective, since f(3)=f(-3)=9, and it therefore has no inverse.

The standard repair is to shrink the domain until injectivity holds. Restrict f(x)=x2 to x0 and it becomes injective, with inverse y. The restriction is a genuine choice: x0 would have worked equally well and would give -y as the inverse. Conventions of this kind, chosen for convenience and then treated as inevitable, are behind the odd-looking domain restrictions on the inverse trigonometric functions two lessons from the end of this course.

Finding an inverse algebraically is a routine. Write y=f(x), solve for x in terms of y using the laws of lesson two, and rename the variables at the end.

Inverses in practice

Example. Find the inverse of f(x)=(2x-1)/(x+3), and state where each function is defined.

Set y=(2x-1)/(x+3) and multiply through by x+3, which is legal since x-3: y(x+3)=2x-1, so xy+3y=2x-1. Gather the x terms on one side: xy-2x=-1-3y, so x(y-2)=-(3y+1) and x=(3y+1)/(2-y). Hence f-1(x)=(3x+1)/(2-x). The domain of f excludes -3 and the domain of f-1 excludes 2, and these correspond: 2 is the value f approaches but never takes, so it is missing from the range of f. Check numerically: f(5)=9/8=1.125, and f-1(1.125)=(3.375+1)/(0.875)=4.375/0.875=5.

Now you. Find the inverse of g(x)=(x+5)/(x-2) and check it at x=7.

Answer

y(x-2)=x+5 gives x(y-1)=2y+5, so g-1(x)=(2x+5)/(x-1). Checking, g(7)=12/5=2.4, and g-1(2.4)=(4.8+5)/1.4=9.8/1.4=7.

The most familiar inverse pair in daily use is temperature conversion. Fahrenheit from Celsius is F(c)=9c/5+32, a function built from a stretch and a shift, and undoing those in reverse order gives C(f)=5(f-32)/9. Check: F(37)=66.6+32=98.6, and C(98.6)=5×66.6/9=37. The two scales agree at exactly one temperature, where c=9c/5+32, which solves to c=-40, a linear equation of the kind the previous lesson handled.

What is still missing

A function has now been defined, restricted, composed and inverted, all algebraically. What has not happened is seeing one. Asking whether f(x)=x3-3x+1 is injective, or what its range is, or how many solutions f(x)=0 has, is hard work by algebra alone and nearly immediate once the function is drawn.

Turning a rule into a picture requires a device for holding an input and its output at once, and that device is the coordinate plane. It arrives in the next lesson, and with it the observation that injectivity, range, and the number of solutions of an equation are all visible at a glance.

The plane and the graph

A function is a rule, and a rule is invisible. Asking whether it ever repeats an output, or how many times it hits zero, or where it is largest, means grinding through algebra for answers that a picture would give away instantly.

The device that supplies the picture is the coordinate plane, and this lesson builds it and puts it to work on the two functions the course has met so far, the linear and the quadratic. It assumes only the algebra of the earlier lessons and the definition of a function.

Two numbers for one point

The idea appears in Descartes' La Géométrie of 1637, published as an appendix to the Discourse on Method, and independently in unpublished work of Fermat. Fix two perpendicular number lines crossing at a point called the origin. Any point in the plane is then named by two numbers, how far along and how far up, written (x,y).

The consequence is larger than the construction. Geometry becomes algebra and algebra becomes geometry: a curve is the set of points whose coordinates satisfy an equation, and an equation is a curve. Two thousand years of Greek geometry, done with compass and straightedge on figures, becomes computation with letters, and problems that needed a clever construction each time fall to routine.

For a function this pairing has a specific use. Plot the point (x,f(x)) for every x in the domain, and the resulting set of points is the graph of f. Everything the rule does is now laid out at once, which is why the rest of this course reasons about functions with one eye on their graphs.

Distance, and the first theorem to fall out

Take two points (x1,y1) and (x2,y2). The horizontal gap between them is x2-x1 and the vertical gap is y2-y1, and those two gaps are the legs of a right triangle whose hypotenuse joins the points. Pythagoras' theorem gives the length immediately:

d=(x2-x1)2+(y2-y1)2

The squares mean the order of subtraction does not matter, which is as it should be, since distance has no direction. The midpoint is even simpler, the average of the coordinates separately, big((x1+x2)/2,(y1+y2)/2big).

Example. Find the distance between (2,-3) and (7,9), and their midpoint.

The gaps are 7-2=5 and 9-(-3)=12, so d=25+144=169=13. The midpoint is (4.5,3).

Now you. Find the distance between (-4,5) and (8,-11).

Answer

The gaps are 12 and -16, so d=144+256=400=20.

Circles come free from the same formula. The set of points at distance r from (a,b) is (x-a)2+(y-b)2=r2, which is a definition turned directly into an equation. Note that a circle is not the graph of a function: the input x=3 has two points above it on a circle centred at the origin, and a function must have exactly one output. This is the vertical line test: a set of points is the graph of a function precisely when no vertical line meets it twice.

The straight line

The defining feature of a line is that the ratio of vertical change to horizontal change is the same everywhere on it. That constant is the slope or gradient m=(y2-y1)/(x2-x1), and it is the rate at which the output changes per unit of input.

Given a slope m and any point (x1,y1) on the line, every other point (x,y) satisfies (y-y1)/(x-x1)=m, which rearranges to the point-slope form y-y1=m(x-x1). Expanding gives y=mx+c, where c is the value at x=0, the intercept. Vertical lines are the exception: their slope is undefined, since the horizontal change is zero, and they are written x=k and are not graphs of functions.

Two lines are parallel when their slopes are equal. They are perpendicular when m1m2=-1, which can be seen by rotating a slope triangle a quarter turn: a rise of a over a run of b becomes a rise of -b over a run of a. So the perpendicular to a line of slope 3 has slope -1/3.

Example. Find the line through (2,7) and (6,19).

The slope is (19-7)/(6-2)=12/4=3. Using the first point, y-7=3(x-2), so y=3x+1. Check the second point: 3×6+1=19.

Now you. Find the line through (-3,11) and (5,-5), and check both points.

Answer

The slope is (-5-11)/(5+3)=-16/8=-2, so y-11=-2(x+3) and y=-2x+5. Checking: -2(-3)+5=11 and -2(5)+5=-5.

A line through real data

Amos Dolbear noticed in 1897 that snowy tree crickets chirp faster when it is warmer, and that the relationship is close to linear over the range where the insects are active. His rule, still quoted, is that the temperature in degrees Fahrenheit is 50 plus a quarter of the amount by which the chirp count per minute exceeds 40.

As a function that is T(N)=50+(N-40)/4, a line of slope 1/4 through the point (40,50). At 120 chirps per minute it predicts 70 °F, at 160 it predicts 80 °F, and at 80 it predicts 60 °F. The slope carries the physical content: four extra chirps per minute for each additional degree, which is a statement about how insect metabolism responds to temperature.

The limits of the model are visible in the same equation, which is the point of writing it down. Set N=0 and it predicts 40 °F, whereas silence in fact means only that the crickets have stopped, and below about 50 °F they do. Extrapolating a fitted line outside the range of the data that produced it is the commonest abuse of a linear model, and the graph makes the abuse obvious: the line continues cheerfully into a region where no data ever existed.

What a graph shows at a glance

Draw a function and its algebraic properties become visual. The domain is the set of x values under which the curve exists; the range is the set of y values it reaches. Solutions of f(x)=0 are the points where the curve crosses the horizontal axis, so the number of solutions can be counted rather than derived. Solutions of f(x)=g(x) are the crossings of two curves, which is why the intersection of y=3x+1 and y=-2x+5 at x=0.8, y=3.4 is both a picture and the solution of a pair of simultaneous equations.

Injectivity, the condition for having an inverse from the previous lesson, becomes the horizontal line test: f is injective exactly when no horizontal line meets its graph twice. The parabola y=x2 fails it at every positive height, which is the visual form of the fact that 3 and -3 share a square.

The graph of an inverse is a reflection of the graph of the function in the line y=x, because reflecting swaps the roles of the coordinates, and swapping input and output is exactly what the inverse does. This is a good check on any inverse computed algebraically.

The parabola, read off its vertex form

The graph of y=ax2+bx+c is a parabola, opening upward when a>0 and downward when a<0. Its most useful feature is its turning point, and completing the square from the third lesson delivers it without further work.

Take y=2x2-12x+5. Factor 2 out of the first two terms: y=2(x2-6x)+5. Complete the square inside: x2-6x=(x-3)2-9, so y=2big[(x-3)2-9big]+5=2(x-3)2-13.

Now everything is visible. The squared term is never negative, so y is smallest when it is zero, at x=3, where y=-13. The vertex is (3,-13), the curve is symmetric about the vertical line x=3, and since the minimum is below zero the parabola crosses the axis twice, at the two roots computed in the earlier lesson, 3±6.5. The discriminant and the vertex are saying the same thing in two languages: the sign of the discriminant is the question of whether the vertex sits on the far side of the axis from the opening.

Example. Put y=x2-6x+5 into vertex form and describe the graph.

x2-6x=(x-3)2-9, so y=(x-3)2-4. The vertex is (3,-4), opening upward, symmetric about x=3. Since y=0 gives (x-3)2=4, the crossings are at x=1 and x=5, and the vertex sits midway between them as symmetry requires.

Now you. Put y=3x2+12x-7 into vertex form and give the vertex.

Answer

y=3(x2+4x)-7=3big[(x+2)2-4big]-7=3(x+2)2-19, so the vertex is (-2,-19). Check by substituting x=-2 into the original: 12-24-7=-19.

Shifting, stretching, reflecting

Vertex form is an instance of something general. Once the graph of y=f(x) is known, four modifications give a family of related graphs, and knowing them removes most of the work of plotting.

Replacing f(x) by f(x)+k moves the graph up by k, which is obvious since every output gains k. Replacing it by f(x-h) moves it right by h, which surprises people, and the reason is that the new function at x=h does what the old one did at x=0: the input must be larger by h to produce the same behaviour. Multiplying the output, af(x), stretches vertically by a factor a, and a negative a also flips the graph upside down. Multiplying the input, f(bx), squeezes horizontally by a factor b, again the opposite of what the symbol suggests, for the same reason as the shift.

So y=2(x-3)2-13 is the basic parabola y=x2 moved three right, stretched by two, and dropped thirteen. Every quadratic is the same curve in a different position, which is a genuinely surprising fact: parabolas differ in scale and place but not in shape. Later lessons apply the same four moves to exponentials and to waves, where a horizontal shift becomes a phase and a horizontal squeeze becomes a frequency.

Where the picture misleads

A graph is evidence, not proof, and the honest limits are worth carrying. A plot shows a window, and behaviour outside it is invisible: y=x3-3x+1 looks like a rising line if plotted from -100 to 100, and only a window a few units wide reveals its two turning points. Features can also be too small to see, and a curve that appears to touch the axis may cross it just below the resolution of the drawing.

Nor does a graph settle exact values. It shows that y=x2-2 crosses the axis somewhere near 1.41; it cannot tell you that the crossing is irrational, which took the argument sketched in the first lesson. Pictures are for finding out what is true. Algebra, and eventually proof, is for establishing it.

Both curves in this lesson are members of one family: the linear function is degree one and the quadratic degree two, and nothing in the definition of a polynomial stops at two. What degree three and beyond look like, and how their roots and their factors are the same information twice, is the next lesson.

Polynomials and their roots

The line and the parabola are the first two members of a family, and there is nothing in either that suggests the family should stop at degree two. Cubics and quartics are written the same way, arise from the same kinds of problem, and are handled by one set of tools.

This lesson builds those tools. It assumes the algebra of the earlier lessons and the graphs of the previous one, and it produces the central fact about polynomials: a root and a factor are the same thing seen twice, so finding one is finding the other.

Degree, and what a polynomial is not

A polynomial in x is a finite sum of terms akxk with whole-number exponents and constant coefficients, such as 3x4-x3+7x-2. The largest exponent with a non-zero coefficient is the degree, and the coefficient sitting on it is the leading coefficient. Degree zero is a constant, one is linear, two quadratic, three cubic, four quartic.

The restriction to whole-number exponents is what makes the family well behaved, and it excludes a great deal. Neither 1/x, which is x-1, nor x, which is x1/2, is a polynomial, and neither is 2x, where the variable is in the exponent rather than the base. Each of those three has its own lesson later in this course, and each behaves quite differently.

Adding two polynomials gives a polynomial whose degree is at most the larger of the two, and can be less when leading terms cancel. Multiplying gives one whose degree is exactly the sum, since the leading terms multiply and nothing else can reach that power. So degree behaves under multiplication the way exponents do, which is the first hint of the index laws that a later lesson makes general.

At large lvertxrvert the leading term swamps everything else, and this settles the shape of the far ends of the graph. For p(x)=x3-4x2+2x+5 at x=100, the leading term is 1{,}000{,}000 while the rest contributes -40{,}000+200+5, four per cent of the total, and the fraction shrinks as x grows. So an odd-degree polynomial with positive leading coefficient falls without limit to the left and rises without limit to the right, and an even-degree one goes the same way at both ends.

Division with a remainder

Whole numbers can be divided with a remainder smaller than the divisor: 17=5×3+2. Polynomials do the same, with the degree of the remainder taking the place of the size: dividing p(x) by a non-zero d(x) gives a unique quotient q(x) and remainder r(x) with

p(x)=d(x)q(x)+r(x)

where the degree of r is less than the degree of d. Dividing by a linear x-a therefore leaves a remainder of degree zero, a plain number.

The procedure is long division, arranged like the arithmetic version. Divide x3-4x2+2x+5 by x-3. The leading term x3 divided by x gives x2; multiply x-3 by x2 to get x3-3x2 and subtract, leaving -x2+2x+5. Now -x2 divided by x gives -x; multiply and subtract, leaving -x+5. Finally -x divided by x gives -1; multiply and subtract, leaving 2. So the quotient is x2-x-1 and the remainder is 2.

Multiplying back is the check, and it should always be done: (x-3)(x2-x-1)=x3-x2-x-3x2+3x+3=x3-4x2+2x+3, and adding the remainder 2 recovers the original.

The remainder theorem, and why it is obvious in hindsight

There is a shortcut for the remainder that avoids the division entirely. Write the division statement for a linear divisor: p(x)=(x-a)q(x)+r, with r a constant. This holds for every x, so it holds at x=a, where the first term dies:

p(a)=0×q(a)+r=r

The remainder on dividing by x-a is simply p(a). Checking against the division just done, p(3)=27-36+6+5=2, which is the remainder found the long way.

The factor theorem is the case r=0, and it is the reason the whole apparatus matters. If p(a)=0 then the remainder is zero, so x-a divides p exactly; and if x-a divides p exactly then p(a)=0. A number a with p(a)=0 is a root, and the theorem says roots and linear factors are two descriptions of one fact. Finding a root hands you a factor, and dividing by that factor leaves a polynomial of one lower degree, which is a smaller problem of the same kind.

Example. Show that 2 is a root of p(x)=2x3+x2-13x+6, and factor it completely.

Evaluate: p(2)=16+4-26+6=0, so x-2 is a factor. Dividing gives 2x2+5x-3, and that quadratic factors as (2x-1)(x+3), since -3×2=-6 and the middle term works out. So p(x)=(x-2)(2x-1)(x+3) and the roots are 2, 1/2 and -3. Check the middle one: p(0.5)=0.25+0.25-6.5+6=0.

Now you. Show that 1 is a root of q(x)=x3-7x+6 and factor it completely.

Answer

q(1)=1-7+6=0, so x-1 is a factor. Dividing gives x2+x-6=(x+3)(x-2), so q(x)=(x-1)(x-2)(x+3) and the roots are 1, 2 and -3. Check: q(2)=8-14+6=0 and q(-3)=-27+21+6=0.

Where to look for a root

The factor theorem is useless without a root to start from, and guessing is only tolerable if the guesses are few. The rational root theorem makes them few. If a polynomial with integer coefficients has a rational root p/q in lowest terms, then p divides the constant term and q divides the leading coefficient.

For 2x3+x2-13x+6 the constant is 6 and the leading coefficient is 2, so the numerator is ±1,±2,±3,±6 and the denominator is 1 or 2. Once repeats such as 2/2 are struck out that is twelve candidates rather than infinitely many, and the roots found above, 2, 1/2 and -3, are all on the list as they must be.

The theorem is a genuine restriction and also a warning about how special rational roots are. Most polynomials with integer coefficients have none: x3-2 has only 23, which is irrational, and the candidates ±1,±2 all fail on inspection. Textbook exercises are chosen to have rational roots, and this creates a false impression that finding roots is a matter of persistence. In the wild it is not, and numerical methods do the work.

Multiplicity, and what the graph does

A factor can repeat. In p(x)=(x-2)2(x+1) the root 2 has multiplicity two, and the graph shows the difference plainly: at a root of odd multiplicity the curve crosses the axis, and at a root of even multiplicity it touches and turns back. The reason is a sign argument. Near x=2 the factor (x-2)2 is positive on both sides, so the sign of the product is controlled by (x+1) alone and does not change; with a single factor (x-2) the sign flips as x passes through.

Counting roots with their multiplicities makes the bookkeeping come out right. A polynomial of degree n has at most n roots, because each root contributes a linear factor and the degrees of the factors must add to n. It may have fewer real roots: x2+1 has none, the gap left open in the first lesson, and x3-x has three while x3+x has one.

The fundamental theorem of algebra, proved by Gauss in his doctoral thesis of 1799 with several later and more rigorous proofs, says that over the complex numbers the count is always exact: every polynomial of degree n1 has exactly n roots counted with multiplicity. This is the payoff of the extension the first lesson declined to make. Admitting the solutions of x2=-1 does not merely patch that one equation, it makes every polynomial equation solvable at once, and no further extension is ever needed. On the real line the statement is weaker and messier, and the honest version is that the roots come in a bounded number and some of them may be missing.

Example. Factor x4-5x2+4 completely and describe its graph.

The expression is a quadratic in x2, so treat u=x2: u2-5u+4=(u-1)(u-4). Substituting back, (x2-1)(x2-4)=(x-1)(x+1)(x-2)(x+2). Four distinct roots, all of multiplicity one, so the curve crosses the axis at -2,-1,1,2. The degree is even with a positive leading coefficient, so both ends rise, and the curve must therefore dip below the axis between -2 and -1, rise above between -1 and 1, and dip again between 1 and 2.

Now you. Factor x3-3x2-4x+12 completely, given that 3 is a root.

Answer

Dividing by x-3 gives x2-4, so the factorisation is (x-3)(x-2)(x+2) and the roots are 3, 2 and -2. Check x=-2 in the original: -8-12+8+12=0.

Building a polynomial from what it must do

The correspondence runs both ways, which is how polynomials get used as models. If a cubic must vanish at -1, 2 and 5, it is a(x+1)(x-2)(x-5) for some constant a, and one further condition fixes a.

Example. Find the cubic with roots -1, 2, 5 passing through (0,20).

Write p(x)=a(x+1)(x-2)(x-5). At x=0 this is a×1×(-2)×(-5)=10a, and that must equal 20, so a=2. Expanding, p(x)=2(x+1)(x2-7x+10)=2(x3-6x2+3x+10)=2x3-12x2+6x+20. Check at x=2: 16-48+12+20=0.

Now you. Find the quadratic with roots 3 and -4 passing through (1,-12).

Answer

p(x)=a(x-3)(x+4), and at x=1 that is a(-2)(5)=-10a=-12, so a=1.2 and p(x)=1.2(x-3)(x+4)=1.2x2+1.2x-14.4.

The limits of exact solution

Everything above is exact, and it depends on someone handing over a root. When no rational root exists, degree three and four can still be solved in radicals by the formulas of Cardano and Ferrari, published in the Ars Magna of 1545, though the cubic formula is unpleasant enough that almost nobody uses it.

Beyond that the road ends. Ruffini argued in 1799 and Abel proved in 1824 that the general fifth-degree equation has no solution by radicals, and Galois in 1832 explained exactly which equations do and which do not, in work that founded modern algebra. This is not a gap awaiting a cleverer formula. It is a proof that no such formula can exist, and it is one of the places where mathematics stops offering better methods and starts explaining why a method is impossible.

What is left is numerical work, which is enough for every practical purpose: bisection, which uses the fact that a sign change traps a root, and faster iterative schemes built on the derivative. Those belong to calculus.

Polynomials share one further limitation, and it is the one that opens the next lesson. No polynomial doubles at fixed intervals. A quantity that grows by a fixed factor for each unit of time, as money at compound interest does and as a bacterial culture does, cannot be described by any of these functions, however high the degree. Describing it needs the variable in the exponent, which is where the family ends and a new one begins.

Exponentials and growth

Money in an account earning five per cent does not grow by a fixed amount each year; it grows by a fixed factor. No polynomial does that, however high its degree, so the family of functions built in the previous lessons cannot describe interest, populations, or radioactive decay.

This lesson builds the family that can. It starts from repeated multiplication, which is arithmetic, and ends with a function defined for every real exponent, a specific irrational number that falls out of compounding, and a comparison that settles which kind of growth wins in the long run.

Counting the factors

Write bn for b multiplied by itself n times. With that reading the index laws are not rules to memorise but observations about counting. Multiplying bm by bn writes down m factors and then n more, so

bm×bn=bm+n

Raising a power to a power writes down n copies of a block of m factors, giving (bm)n=bmn. Distributing across a product, (ab)n=anbn, because multiplication is commutative and the factors can be sorted. That is the whole of the theory for whole-number exponents, and every one of the three laws is proved by counting.

Division follows from the first law: bm/bn=bm-n whenever m exceeds n, since cancelling n factors from m leaves m-n. So 25×23/26=22=4, which is worth checking directly: 32×8/64=4.

The extensions are forced, not chosen

Nothing in the counting picture makes sense of b0, or b-3, or b1/2: multiplying something by itself zero times, or minus three times, is not an operation. Yet all three have standard values, and the reason is a principle used repeatedly in the first lesson. Extend the notation in whatever way keeps the existing laws true, then check that nothing already working breaks.

Apply the first law with n=0: bm×b0=bm+0=bm. Dividing by bm, which is legal when b is not zero, gives b0=1. That is not a convention adopted for tidiness, it is the only value compatible with the law. Now apply the law with n=-m: bm×b-m=b0=1, so b-m must be 1/bm. A negative exponent is a reciprocal, forced.

Fractional exponents come from the second law. Whatever b1/2 is, squaring it gives (b1/2)2=b1=b, so it is a square root of b, and the positive root is chosen to keep the function single valued for positive b. In general bp/q is the q-th root of bp. So 82/3 is the cube root of 64, which is 4, and equally the square of the cube root of 8, which is 22=4. Both routes must agree, and they do.

Two restrictions come with this. The base is kept positive, because (-8)1/2 has no real value while (-8)1/3=-2 does, and a function that exists at some fractional exponents and not others is unusable. And 00 is left undefined, since the pattern b0=1 and the pattern 0n=0 give different answers and neither has priority.

Example. Evaluate 16-3/4 and simplify (2a3)4.

For the first, the negative sign inverts and the fraction takes a root: 16-3/4=1/163/4, and 163/4 is the cube of the fourth root of 16, which is 23=8. So the value is 1/8=0.125. For the second, distribute the exponent across the product: 24(a3)4=16a12. Check at a=2: the left is (2×8)4=164=65{,}536 and the right is 16×4096=65{,}536.

Now you. Evaluate 274/3 and simplify (3x2)3x-4.

Answer

274/3 is the fourth power of the cube root of 27, which is 34=81. And (3x2)3x-4=27x6x-4=27x2. Check at x=2: 27×64×(1/16)=108, and 27×4=108.

Irrational exponents, and where the real line earns its keep

Every exponent so far has been rational. What is 22?

The honest answer needs the completeness of the real line from the first lesson. The rationals approach 2 as closely as you like, and the corresponding powers approach a single value: 21.4=2.639016, 21.41=2.657372, 21.414=2.664750, and the sequence closes in on 2.665144. Completeness guarantees that a real number sits exactly where the sequence is heading, and that number is defined to be 22.

The definition is worth pausing on, because it is the first place in this course where a value is defined by a limiting process rather than computed by an operation. The gap the Pythagoreans found is what makes the definition necessary, and the completeness Dedekind and Cantor supplied is what makes it work. With that step taken, bx is defined for every real x and every positive base b, the index laws still hold, and the result is a genuine function with domain all of R.

Its range is only the positive numbers. Raising a positive base to any real power never produces zero or a negative, which means the graph of y=2x lies entirely above the horizontal axis, approaching it as x falls without ever touching. For b>1 the function increases and models growth; for 0<b<1 it decreases and models decay; and b=1 gives the constant 1, which is why that base is excluded from anything interesting.

Compound interest, and the number it converges on

Put P in an account paying an annual rate r, compounded n times a year. Each period multiplies the balance by 1+r/n, and after t years there have been nt periods, so

A=P(1+rn)nt

Example. Put 1000 pounds at 6 per cent for five years, compounded monthly. What is the balance?

Here r/n=0.06/12=0.005 and nt=60, so A=1000×1.00560=1000×1.34885=1348.85 pounds. Compounding annually instead gives 1000×1.065=1338.23, and daily gives 1349.83. More frequent compounding pays more, and the increments are shrinking.

Now you. Put 2500 pounds at 4.5 per cent for eight years, compounded monthly.

Answer

r/n=0.00375 and nt=96, so A=2500×1.0037596=2500×1.432365=3580.91 pounds.

The shrinking increments raise the obvious question: what happens as the compounding gets infinitely frequent? Take one pound at one hundred per cent for one year, so the balance is (1+1/n)n, and push n up. Annually it gives 2. Semi-annually, 2.25. Monthly, 2.613035. Daily, 2.714567. Hourly, 2.718127. At a million periods, 2.718280.

The sequence does not run away; it converges, and its limit is

e=2.718281828459045dots

named by Euler, who computed it to eighteen places in 1748 and proved it irrational. Continuous compounding therefore gives A=Pert, and for the example above that is 1000×e0.3=1349.86 pounds, three pence more than daily compounding. The interesting part is not the three pence. It is that a limit of purely financial bookkeeping produces a constant that turns out to be the natural base for every growth process, for reasons that only calculus makes fully clear: ex is the unique exponential whose rate of increase at each point equals its own value there.

Decay, and half-lives

When the base is less than one the same function runs downhill. Radioactive decay is the cleanest case, because the fraction lost per unit time is fixed by physics and unaffected by temperature, pressure or chemistry. It is usual to write the law using the half-life T, the time for half of any sample to decay:

N(t)=N0(12)t/T

Carbon-14 has a half-life of 5730 years. After three half-lives, that is 17{,}190 years, an eighth is left, which needs no calculation. For a time that is not a whole number of half-lives, the fractional exponent does the work.

Example. What fraction of the carbon-14 in a sample remains after 15{,}000 years?

The exponent is 15000/5730=2.6178, so the fraction is 0.52.6178=0.1629, about 16 per cent. As a sanity check, this sits between the eighth left after three half-lives and the quarter left after two, nearer the eighth, which matches an exponent nearer three than two.

Now you. Caesium-137, the isotope of most concern after the Chernobyl accident, has a half-life of 30.08 years. What fraction remains after a century?

Answer

The exponent is 100/30.08=3.324, so the fraction is 0.53.324=0.0998, just under a tenth.

Why exponentials always win

Set an exponential against a polynomial and the exponential eventually wins, no matter how modest its base or how high the polynomial's degree. The comparison is worth doing with real numbers, because the crossover can be a long way out and the intermediate behaviour is genuinely misleading.

Compare 2x with x3. At x=5 the cube is well ahead, 125 against 32. At x=9 it is still ahead, 729 against 512. At x=10 the exponential has taken the lead, 1024 against 1000, and it never loses it again: at x=20 the figures are 1{,}048{,}576 against 8000.

Raise the stakes to x10 and the exponential looks hopeless for a long time, trailing by many orders of magnitude through the whole of the range anyone would plot. It overtakes just before x=59, where 259=5.76×1017 against 5910=5.11×1017, and thereafter the gap widens without limit. The pattern is general: any exponential with base above one eventually exceeds any polynomial, permanently. This single fact is why an algorithm whose cost grows exponentially with the size of its input is useless at scale while a cubic one is merely slow.

What nothing grows exponentially forever

Exponential models describe the early part of many processes and the whole of almost none, and saying so is not a hedge. A single bacterium dividing every twenty minutes would after eight hours be 224 cells, about seventeen million, which is realistic in fresh broth. Continue for two days and the same model predicts a mass exceeding that of the Earth. What stops it is not the mathematics but food, space and waste, and the curve bends over into a shape called logistic that this course does not cover.

The same warning applies to money and to populations. Compound interest is exact as long as the rate holds, and rates do not hold for centuries. World population grew at over two per cent a year in the late 1960s and grows at under one per cent now, so any exponential fitted then would badly overshoot today.

Use exponentials for what they are: the correct description of a fixed proportional rate, valid while that rate is fixed. Within that scope they are exceptionally accurate, which raises the practical question this lesson cannot answer. Computing 2x is easy for any x. Going backwards, finding the x for which 2x=10, has no method here at all. That inverse function is the next lesson.

Logarithms

Computing 2x for any x is straightforward, and the previous lesson did it repeatedly. Going the other way, finding the x for which 2x=10, has no method at all in anything covered so far, and it is the question every growth problem eventually asks: not what the balance will be, but when it doubles.

This lesson defines the function that answers it, derives its three laws from the index laws, and puts it to work on dating, doubling and the measurement scales that span too many orders of magnitude to plot any other way.

The inverse of an exponential

For a fixed base b greater than zero and not equal to one, the exponential y=bx passes the horizontal line test: it increases steadily when b>1 and decreases steadily when b<1, so no output is ever repeated. By the criterion of the fourth lesson it therefore has an inverse, and that inverse is the logarithm to base b:

logby=xmeans exactlybx=y

Read it as the question "what power of b gives y", and most confusion about logarithms disappears. So log28=3 because 23=8; log101000=3; log51=0, since any base to the power zero is one.

Being an inverse fixes its domain and range without further argument. The range of bx is the positive reals, so that is the domain of the logarithm: there is no logarithm of zero or of a negative number, because no power of a positive base produces one. The domain of bx is all of R, so that is the range: a logarithm can be any real number, positive or negative. And by the reflection rule of the fifth lesson, the graph of y=logbx is the graph of y=bx flipped in the line y=x, rising steeply near zero, crossing the axis at x=1, and thereafter climbing ever more slowly.

Two bases are so common they have their own notation. Base 10 is the common logarithm, written log, and base e is the natural logarithm, written ln. Base 2 appears throughout computing, where it counts bits.

Three laws, and where they come from

Each index law from the previous lesson becomes a logarithm law when read backwards. Let M=bm and N=bn, so that m=logbM and n=logbN. Then MN=bmbn=bm+n, and taking the logarithm of both sides gives

logb(MN)=m+n=logbM+logbN

The logarithm turns multiplication into addition. The same move on M/N=bm-n gives logb(M/N)=logbM-logbN, and on Mk=(bm)k=bmk gives

logb(Mk)=klogbM

That third law is the one that solves equations, because it takes an unknown out of an exponent and puts it in front where algebra can reach it.

A fourth relation lets any base be converted to any other, which matters because calculators carry only two. If x=logby then bx=y; take natural logarithms of both sides and use the third law to get xlnb=lny, so

logby=lnylnb

The choice of ln here is arbitrary and base 10 works identically, which is a useful check: log3200 computed as ln200/ln3=5.2983/1.0986=4.8227, and as log200/log3=2.3010/0.4771=4.8227.

Two errors are worth naming because they are so common. There is no law for log(M+N), which does not simplify at all. And log(M)/log(N) is not log(M/N); the first is a change of base and the second is a difference.

Solving for an exponent

The routine has three steps: isolate the exponential, take logarithms of both sides, and use the third law to bring the exponent down.

Example. Solve 3x=200.

Take natural logarithms: ln(3x)=ln200, so xln3=ln200 and x=ln200/ln3=5.2983/1.0986=4.8227. Check: 34.8227=199.99, which is the intended 200 to the precision of the rounded exponent.

Now you. Solve 5x=1000.

Answer

x=ln1000/ln5=6.9078/1.6094=4.2920. Check: 54.292=999.95.

Doubling time is this calculation with the answer 2 fixed in advance. Money at an annual rate r multiplies by (1+r)t, so doubling needs (1+r)t=2, giving t=ln2/ln(1+r). At three per cent that is 0.6931/0.029559=23.45 years; at five per cent, 14.21 years; at eight per cent, 9.01 years.

The familiar rule of thumb, that dividing 72 by the percentage rate gives the doubling time, is now visible as an approximation to this formula: it predicts 24, 14.4 and 9 against the true 23.45, 14.21 and 9.01. It is good to within a few per cent across the range of rates anyone encounters, and it is wrong for very large rates, where the logarithm's curvature bites.

Dating the past

Radiocarbon dating is the same equation solved for time rather than for a rate. Living tissue holds carbon-14 at the atmospheric ratio; from death the isotope decays with a half-life of 5730 years and is not replenished, so measuring what fraction remains gives an age. Willard Libby developed the method in 1949 and received the Nobel Prize for chemistry in 1960 for it.

The decay law from the previous lesson is N/N0=(1/2)t/T. Take logarithms of both sides and use the third law:

t=Tln(N/N0)ln(1/2)

Example. A piece of charcoal retains 22 per cent of the carbon-14 of living wood. How old is it?

Substituting, t=5730×ln(0.22)/ln(0.5)=5730×(-1.5141)/(-0.6931)=5730×2.1844=12{,}517 years. A check on plausibility: two half-lives would leave 25 per cent at 11{,}460 years, and 22 per cent is a little less, so the age must be a little more, as it is.

Now you. A bone retains 35 per cent. How old is it?

Answer

t=5730×ln(0.35)/ln(0.5)=5730×1.5146=8679 years.

The technique has honest limits, and practitioners state them. Beyond about 50{,}000 years, roughly nine half-lives, too little carbon-14 remains to measure against background. The atmospheric ratio is not in fact constant, so raw dates are corrected against tree-ring and coral records, a calibration that can move a result by centuries. And nuclear weapons testing in the 1950s nearly doubled atmospheric carbon-14, which makes recent material behave strangely and, incidentally, allows the year of formation of tissue from that era to be dated precisely.

Why the logarithm was a technology

Logarithms were invented as a labour-saving device, not as a piece of theory. John Napier published Mirifici Logarithmorum Canonis Descriptio in 1614 after twenty years of computation, and Henry Briggs recast the idea to base 10 and published tables to fourteen places in 1624. The point was the first law. Before mechanical calculators, multiplying two eight-digit numbers by hand was slow and error-prone, while adding them was easy, so a table that converted every multiplication into an addition roughly halved the labour of astronomical calculation. Laplace said it doubled the life of the astronomer.

The slide rule is that idea in wood. Two rulers marked not in units but in the logarithms of units, slid against each other, add lengths that represent logarithms and therefore multiply the numbers. Every engineer carried one until pocket calculators arrived in the early 1970s, and the Apollo programme was flown with them.

Two habits of that era survive because they remain useful. The characteristic, the whole-number part of a base-10 logarithm, is the order of magnitude: log(6.02×1023)=23.78, so Avogadro's number has 24 digits. And in computing, log2 counts how many times a quantity can be halved, which is why searching a sorted list of a trillion items takes at most log2(1012)=39.9, that is 40, comparisons.

Scales that span too much to plot

When a quantity ranges over many orders of magnitude, plotting it linearly wastes almost the whole axis on the top end. Taking logarithms first compresses it into something a person can read, and several measurement scales have this built in.

Acidity is defined as pH=-log10[H+], with the concentration in moles per litre. Sound level in decibels is 10log10(I/I0) against a reference intensity, so an increase of 30 dB is a factor of 103=1000 in intensity: ordinary conversation at 55 dB carries a thousandth of the energy of a food blender at 85 dB. Earthquake magnitude, as defined by Charles Richter in 1935, adds one unit per factor of ten in ground-motion amplitude, which corresponds to a factor of about 31.6 in energy, since energy scales as the 1.5 power of amplitude. Stellar magnitudes, formalised by Norman Pogson in 1856 to match an ancient naked-eye scale, put five magnitudes to a factor of 100 in brightness, so one magnitude is a factor of 1001/5=2.512, and the scale runs backwards with brighter objects having smaller numbers.

Example. A solution has a hydrogen ion concentration of 3.2×10-4 mol per litre. What is its pH, and how does it compare with pure water at pH 7?

pH=-log10(3.2×10-4)=-(log103.2-4)=4-0.5051=3.495. That is 3.5 units below neutral, so the concentration is about 103.5=3160 times that of pure water. This is roughly the acidity of orange juice.

Now you. Surface seawater has a hydrogen ion concentration near 7.9×10-9 mol per litre. What is its pH?

Answer

pH=-log10(7.9×10-9)=9-log107.9=9-0.8976=8.102, mildly alkaline.

How slowly it grows

The previous lesson showed the exponential outrunning every polynomial. Reflecting that statement in the line y=x gives its counterpart: the logarithm grows more slowly than every positive power of x, however small the power. It does grow without limit, since log2x reaches 100 when x reaches 2100, but it takes its time.

That slowness is exactly why logarithmic scales work. Multiplying the input by a fixed factor adds a fixed amount to the output, so a plot of the logarithm turns proportional change into equal steps, and data spanning a factor of a billion fits on one axis.

Growth and decay are now fully covered: exponentials go forward, logarithms come back, and each is the other's mirror. What none of this describes is anything that repeats. A tide, a note, a rotating shaft and an alternating current return to the same value again and again, and no polynomial, exponential or logarithm can do that, since each is either injective or eventually monotone. Functions that repeat come from angles, and the next lesson starts with the triangle.

Trigonometry of the right triangle

Nothing built so far repeats. Polynomials, exponentials and logarithms all eventually head off in one direction and stay there, and yet a tide, a note and a rotating shaft all come back to where they started, over and over.

The functions that repeat come from angles, and this lesson builds them where they were first built, inside a right triangle. It assumes the algebra of the earlier lessons and Pythagoras' theorem, and it ends with two rules that solve any triangle at all, including those with no right angle in them.

One angle fixes every ratio

Take a right triangle with one acute angle θ. Now draw a bigger one with the same angle θ. The two triangles have the same three angles, since the right angles match and the third is whatever is left of 180°, so they are similar: one is a scaled copy of the other.

Similar triangles have proportional sides. Doubling every side of the first triangle gives the second, or tripling, or multiplying by 1.7, and in every case the ratio of any two sides within a triangle is unchanged, because both members of the ratio were multiplied by the same factor. This is the fact everything below rests on: the ratios depend on the angle alone, not on the size of the triangle.

So the ratios can be named as functions of the angle. Label the side opposite θ as opposite, the one next to it that is not the hypotenuse as adjacent, and define

sinθ=oppositehypotenuse,cosθ=adjacenthypotenuse,tanθ=oppositeadjacent

Dividing the first by the second cancels the hypotenuse and leaves the third, so tanθ=sinθ/cosθ always, which is one fewer thing to remember. Because the hypotenuse is the longest side, sine and cosine of an acute angle always lie strictly between 0 and 1, while the tangent is unbounded and grows without limit as θ approaches 90°.

Two triangles worth knowing exactly

Most values of these functions are irrational and are looked up or computed. Two triangles give exact values, and they are worth knowing because they appear constantly.

Cut a square of side 1 along its diagonal. The result is a right triangle with two legs of 1 and, by Pythagoras, a hypotenuse of 2, and its acute angles are both 45°. So sin45°=cos45°=1/2=0.707107 and tan45°=1.

Cut an equilateral triangle of side 2 down the middle. The result is a right triangle with hypotenuse 2, short leg 1, and remaining leg 4-1=3, with angles of 30° and 60°. Reading the ratios off gives sin30°=1/2, cos30°=3/2=0.866025, tan30°=1/3=0.577350, and the same three with sine and cosine swapped at 60°. The swap is not a coincidence: the two acute angles of a right triangle add to 90°, and one angle's opposite side is the other's adjacent, so sinθ=cos(90°-θ) for every acute θ. That relation is where the word cosine comes from, the sine of the complement.

Pythagoras' theorem itself becomes an identity in this language. In a right triangle with hypotenuse 1, the legs are exactly cosθ and sinθ, so a2+b2=c2 reads

sin2θ+cos2θ=1

Checking at 37°: 0.601822+0.798642=0.36219+0.63781=1. This is the most used identity in the subject, and it is Pythagoras wearing different clothes.

Solving a right triangle

Given one side and one acute angle, the other two sides follow. The discipline is to write the ratio that connects what you know to what you want, then rearrange.

Example. Standing 24 m from the base of a tree, an observer whose eye is 1.6 m above the ground measures the angle of elevation of the top as 37°. How tall is the tree?

The horizontal distance is adjacent to the angle and the height above eye level is opposite, so the tangent connects them: tan37°=h/24, giving h=24×0.75355=18.09 m above eye level. The tree is 18.09+1.6=19.69 m tall. The eye height matters: omitting it understates the tree by more than eight per cent.

Now you. A ladder 5.2 m long leans against a wall at 68° to the horizontal. How far up the wall does it reach, and how far is its foot from the wall?

Answer

The reach is opposite the angle, so it is 5.2sin68°=5.2×0.92718=4.821 m, and the foot is 5.2cos68°=5.2×0.37461=1.948 m from the wall.

Going the other way, from a ratio to the angle, needs the inverse functions sin-1, cos-1 and tan-1, also written arcsin, arccos and arctan. Here the restriction from the fourth lesson bites. Sine is not injective over all angles, so it has no inverse until its domain is cut down, and the convention takes sin-1 to return an angle between -90° and 90°. That is a choice, and it is the reason a calculator answers sin-1(0.5) with 30° and never mentions 150°, which has the same sine. Anyone solving a real problem must decide whether the other angle is the one wanted.

A road signed as a gradient of 1 in 8 rises one metre for every eight along, so its angle is tan-1(0.125)=7.13°. Signs quoted as percentages mean the same thing: a 12 per cent grade is tan-1(0.12)=6.84°. Road gradients are always small enough that the angle in radians, the tangent and the sine agree to within a fraction of a per cent, which is why the sloppiness rarely matters in practice and matters completely on a roof or a ramp.

Reaching triangles with no right angle

Real triangles are rarely right angled, and the definitions above do not apply to them directly. They can be made to apply by dropping a perpendicular, and doing so produces two rules that cover every case.

Label a triangle with angles A, B, C and the side opposite each with the matching lower-case letter. Drop the altitude h from C to the side c. It splits the triangle into two right triangles, and in each of them h is opposite a known angle: h=bsinA from one, h=asinB from the other. Setting these equal gives bsinA=asinB, that is

asinA=bsinB=csinC

the third ratio following by dropping a different altitude. This is the sine rule, and it solves a triangle whenever an angle and its opposite side are both known.

Example. A triangle has A=41°, B=63° and a=12. Find the remaining angle and sides.

The angles sum to 180°, so C=76°. Then b=asinB/sinA=12×0.89101/0.65606=16.297 and c=asinC/sinA=12×0.97030/0.65606=17.748. A check on plausibility: the largest side should face the largest angle, and it does.

Now you. A triangle has A=38°, B=57° and a=9. Find C, b and c.

Answer

C=85°. Then b=9×0.83867/0.61566=12.26 and c=9×0.99619/0.61566=14.563.

The sine rule has a trap that must be stated. Given two sides and an angle not between them, the rule can produce two valid triangles. With a=7, b=10 and A=40°, it gives sinB=10sin40°/7=0.9183, and both B=66.67° and B=113.33° satisfy that, since supplementary angles have equal sines. Each leads to a genuine triangle, one with C=73.31° and one with C=26.69°. This is the ambiguous case, and no amount of algebra resolves it: the data really do describe two shapes, and only extra information settles which.

The cosine rule

When no angle and its opposite side are both known, the sine rule cannot start. Two cases fall outside it: three sides given, or two sides and the angle between them. The cosine rule covers both:

c2=a2+b2-2abcosC

Its derivation is the same altitude trick. Drop the perpendicular from B to the side b, splitting it into pieces of length acosC and b-acosC, with height asinC. Pythagoras on the right-hand triangle gives c2=(asinC)2+(b-acosC)2, and expanding gives a2sin2C+b2-2abcosC+a2cos2C. The two squared trigonometric terms combine to a2 by the Pythagorean identity, leaving the rule.

Set C=90° and the cosine term vanishes, recovering c2=a2+b2. So Pythagoras is the special case, and the -2abcosC is the correction for the angle being something other than a right angle.

Example. Two sides of a triangle are 7 and 9 with an included angle of 52°. Find the third side and the area.

c2=49+81-2×63×cos52°=130-126×0.61566=130-77.57=52.43, so c=7.241. The area of any triangle is half the product of two sides and the sine of the angle between them, which follows from taking one side as the base and asinC as the height: area =0.5×63×sin52°=0.5×63×0.78801=24.82.

Now you. Two sides are 11 and 6 with an included angle of 78°. Find the third side and the area.

Answer

c2=121+36-132cos78°=157-132×0.20791=129.56, so c=11.382. The area is 0.5×66×sin78°=0.5×66×0.97815=32.28.

Rearranged, the cosine rule finds an angle from three sides: cosC=(a2+b2-c2)/2ab. For sides 5, 7 and 10 the angle opposite the longest is cos-1big((25+49-100)/70big)=cos-1(-0.3714)=111.8°, and the negative cosine correctly signals an obtuse angle. Unlike the sine rule, this version is never ambiguous, because cosine takes each value once between 0° and 180°.

Where these definitions stop

Everything in this lesson has been about an angle inside a right triangle, which means an angle strictly between 0° and 90°. Ask for sin150° and the definitions give nothing: no right triangle has an angle of 150° in it, so there is no opposite side to divide by a hypotenuse.

Yet the ambiguous case just used sin113.33°, and the cosine rule just used cos111.8°, both of which were treated as ordinary numbers. Something has been assumed that has not been defined, and that is a debt to settle rather than a detail to skip.

The repair is to leave the triangle behind and put the angle at the centre of a circle, where it can be as large as you like and can turn as many times as it likes. That construction gives sine and cosine a value for every real number, produces the repeating functions this course still owes, and is the next lesson.

The unit circle and periodic functions

The previous lesson used cos111.8° without ever saying what it meant, because a right triangle has no angle larger than 90° and the definitions given there therefore say nothing about one. That debt has to be paid before any of it can be trusted.

Paying it means abandoning the triangle as the home of the definitions and putting the angle at the centre of a circle instead. The reward is much larger than the repair: sine and cosine acquire a value for every real number, and they become the first functions in this course that repeat.

Putting the angle at the centre

Draw a circle of radius 1 centred on the origin. Measure an angle θ anticlockwise from the positive horizontal axis, and let P be the point where that ray meets the circle. Define

cosθ=the x coordinate of P,sinθ=the y coordinate of P

and tanθ=sinθ/cosθ as before.

For an acute angle this agrees with the old definition rather than replacing it. Drop a perpendicular from P to the horizontal axis and you have a right triangle with hypotenuse 1, adjacent side x and opposite side y, so the old ratios give exactly cosθ=x/1 and sinθ=y/1. Nothing already working has broken, which is the same test applied at every extension in this course.

What the new definition adds is that nothing stops θ. Keep turning past 90° and the point keeps moving round the circle, so every angle up to 360° has coordinates, and beyond that the point simply laps: θ and θ+360° give the same point, hence the same values. Negative angles turn clockwise. Sine and cosine are now defined for every real number, and they are periodic with period 360°, meaning sin(θ+360°)=sinθ for all θ.

Since P is on the unit circle, its coordinates satisfy x2+y2=1, which is the identity sin2θ+cos2θ=1 again, now valid for every angle rather than acute ones only. The equation of the circle and the identity are the same statement.

Signs, and values all the way round

The coordinates carry signs, so the functions do too. In the first quadrant both are positive. In the second, from 90° to 180°, the point is left of the axis and above it, so cosine is negative and sine positive. In the third both are negative, and in the fourth cosine is positive and sine negative. The tangent, being the ratio, is positive where the two agree in sign, so it is positive in the first and third quadrants.

Values follow from the acute ones by reflection. The point at 150° is the mirror image of the point at 30° across the vertical axis, so it has the same height and the opposite horizontal position: sin150°=0.5 and cos150°=-0.866025. That is the general rule, and it is why the previous lesson's ambiguous case had two angles with the same sine.

The quadrantal angles complete the picture. At 0° the point is (1,0), at 90° it is (0,1), at 180° it is (-1,0), and at 270° it is (0,-1). So cos90°=0, and tan90° does not exist, because it would require dividing by that zero. The tangent has these breaks every 180°, and its own period is 180° rather than 360°, since diametrically opposite points have coordinates of opposite sign whose ratio is unchanged.

Example. Find sin210°, cos210° and tan210° exactly.

The angle is 30° past 180°, so the point is the reflection of the 30° point through the origin, in the third quadrant where both coordinates are negative. Hence sin210°=-1/2 and cos210°=-3/2=-0.866025, and the tangent is their ratio, 1/3=0.577350, positive as the third quadrant requires.

Now you. Find sin300°, cos300° and tan300° exactly.

Answer

300° is 60° short of a full turn, so the point is the 60° point reflected below the axis: sin300°=-3/2=-0.866025, cos300°=1/2, and tan300°=-3=-1.732051.

Radians, and why degrees are the odd choice

Degrees are inherited from Babylonian astronomy and 360 is arbitrary, chosen because it is close to the number of days in a year and divides conveniently. A unit with a mathematical reason behind it measures the angle by the arc it cuts.

Define one radian as the angle subtending an arc equal in length to the radius. A full circle has circumference 2πr, so a full turn is 2π radians, and

180°=π radians,1 radian=57.296°

Converting is multiplication by π/180 one way and 180/π the other. So 150° is 150π/180=5π/6=2.618 radians, and 210° is 7π/6=3.665 radians.

The payoff is that formulas lose their conversion factors. An arc of angle θ radians on a circle of radius r has length s=rθ, and the sector it bounds has area A=12r2θ, both of which acquire ugly factors of π/180 if degrees are used. The deeper reason, which belongs to calculus, is that the derivative of sinx is cosx only when x is in radians; in degrees an extra factor of π/180 appears and never goes away. From here on, angles are in radians unless a degree sign says otherwise.

Example. A circular sector has radius 2.5 m and angle 1.2 radians. Find its arc length and area.

The arc is s=2.5×1.2=3 m. The area is A=0.5×2.52×1.2=3.75 m². As a check, 1.2 radians is a little under a fifth of a full turn, and a fifth of the full circle's area π×6.25=19.63 m² would be 3.93 m², which is close and slightly larger as expected.

Now you. A sector has radius 6.4 m and angle 0.85 radians. Find its arc length and area.

Answer

s=6.4×0.85=5.44 m and A=0.5×40.96×0.85=17.408 m².

The graphs, and the vocabulary of waves

Plot y=sinx against x in radians and the height of the circling point traces a smooth wave: zero at 0, up to 1 at π/2, back to zero at π, down to -1 at 3π/2, and home at 2π, then repeating forever. Cosine is the same curve shifted left by π/2, since the horizontal coordinate leads the vertical by a quarter turn. That single observation is the identity cosx=sin(x+π/2).

The tangent looks nothing like either. It runs from minus infinity to plus infinity between consecutive breaks at odd multiples of π/2, passing through zero at every multiple of π, and repeats with period π.

The transformations of the fifth lesson now acquire names. In

y=Asinbig(k(x-d)big)+c

lvertArvert is the amplitude, half the distance from trough to peak; c is the midline the wave oscillates about; k compresses the horizontal axis, so the period is 2π/k; and d is the phase shift, moving the whole wave right by d. Any quantity that oscillates smoothly between two extremes can be fitted by choosing these four numbers, which is what makes the sine function a modelling tool rather than a geometrical curiosity.

Fitting a wave to real data

Daylight length is periodic with a period of a year. In London the longest day, around 21 June, has about 16 hours 38 minutes of daylight, and the shortest, around 21 December, about 7 hours 50 minutes.

Those two numbers fix two of the four parameters. The midline is the average, (16.633+7.833)/2=12.233 hours, and the amplitude is half the difference, (16.633-7.833)/2=4.4 hours. The period is 365 days, so k=2π/365. The phase is set by the equinox around 21 March, day 80 of the year, where the length is at its midline and rising. So with t the day of the year,

D(t)=12.233+4.4sin(2π(t-80)365)

Test it away from the points used to build it. On 1 January, t=1, the model gives 7.93 hours, that is 7 hours 56 minutes, against an actual figure near 7 hours 54 minutes. That is close, and the residual error is real rather than rounding: the true curve is not exactly sinusoidal, because the Earth's orbit is an ellipse and its speed varies through the year. A model this simple getting within a few minutes across a whole season is a fair result, and claiming better would be dishonest.

Identities the circle hands over

Several relations that look like formulas to memorise are visible facts about the circle. Reflecting the point across the horizontal axis sends θ to -θ, keeping the horizontal coordinate and flipping the vertical, so cos(-θ)=cosθ and sin(-θ)=-sinθ: cosine is even, sine is odd. Reflecting across the vertical axis sends θ to π-θ and gives sin(π-θ)=sinθ, the supplementary-angle fact that caused the ambiguous case in the previous lesson.

Two identities do not fall out by inspection and have to be proved, most cleanly by rotating a pair of points and computing a distance:

sin(A+B)=sinAcosB+cosAsinB,cos(A+B)=cosAcosB-sinAsinB

They are worth checking numerically rather than taking on trust. With A=45° and B=30°, the first gives sin75°=0.707107×0.866025+0.707107×0.5=0.965926, and a calculator agrees to six places. Notice the minus sign in the cosine formula, which is the commonest error in the subject, and notice that sin(A+B) is emphatically not sinA+sinB, which would give 1.207 here.

Setting B=A gives the double-angle forms sin2A=2sinAcosA and cos2A=cos2A-sin2A, and combining the second with the Pythagorean identity gives cos2A=1-2sin2A, which is the form used to integrate a squared sine in calculus.

Solving equations that have infinitely many answers

An equation such as 2sinx=1 has no single solution, because the sine takes every value in its range once per lap and there are infinitely many laps. The routine is to find the solutions in one turn and then add the period.

Example. Solve 2sinx=1 for 0x<2π.

Divide to get sinx=0.5. One angle with that sine is π/6=0.5236. The other point at the same height is its mirror image in the vertical axis, at π-π/6=5π/6=2.618. Both are in range, so those are the two solutions, and the full solution set over all real x is these two plus any multiple of 2π.

Now you. Solve 2cosx=-2 for 0x<2π.

Answer

cosx=-0.707107, so the reference angle is π/4 and cosine is negative in the second and third quadrants. The solutions are x=π-π/4=3π/4=2.356 and x=π+π/4=5π/4=3.927.

Two habits prevent most errors here. Check which quadrants the sign allows before computing anything, since a calculator's inverse function returns only one of the two. And check the interval you were asked for, because an answer outside it is wrong however correct the arithmetic.

What comes next

The sine function has a property none of the earlier functions had: its values follow a rule that carries from one step to the next, so that knowing sinθ and cosθ gives sin(θ+h) for a fixed step h without starting over. A quantity generated step by step from the previous one, rather than computed directly from its position, is a different kind of object.

Such step-by-step patterns, their closed forms, and what happens when infinitely many of their terms are added together are the subject of the next lesson.

Sequences and series

Some quantities are not computed from where they are but from what came before: each month's balance from last month's, each term from the one in front of it. A rule of that kind generates a list rather than a curve, and it needs its own machinery.

This lesson builds that machinery for the two patterns that cover most of the useful cases, and then asks the question that makes the subject interesting: what it can mean to add infinitely many numbers and get a finite answer. It assumes the exponentials of the earlier lessons and nothing else.

Sequences, explicit and recursive

A sequence is a function whose domain is the natural numbers, written a1,a2,a3,dots rather than a(1),a(2), which is notation rather than a new idea. There are two ways to specify one, and the difference matters in practice.

An explicit rule gives the n-th term directly, so an=3n+2 produces 5,8,11,14 and answers "what is the thousandth term" in one step. A recursive rule gives each term from its predecessors, so a1=5 with an+1=an+3 produces the same list but reaches the thousandth term only by generating the nine hundred and ninety-nine before it.

Recursive definitions are how processes actually work: a bank balance, a population census, a loan. Explicit rules are what you want for computing. Much of this lesson consists of turning a recursive description into an explicit formula, which is called finding a closed form, and it converts a thousand steps of arithmetic into one.

Arithmetic sequences and the pairing trick

A sequence is arithmetic when consecutive terms differ by a constant d, the common difference. Starting from a1, each step adds d, so after n-1 steps

an=a1+(n-1)d

The count n-1 rather than n is where most errors live, and the check is to set n=1 and confirm the formula returns a1.

Adding the terms gives an arithmetic series, and the trick for summing it is the one Gauss is said to have found as a schoolboy asked to add the numbers from 1 to 100. Write the sum forwards and backwards, one under the other, and add columnwise. Each column gives the same total, a1+an, and there are n columns, so twice the sum is n(a1+an) and

Sn=n(a1+an)2

For 1 to 100 that is 100×101/2=5050. The formula reads as the number of terms times their average, which is a useful way to remember it and is exactly what the pairing shows.

Example. A sequence starts at 7 and rises by 4 each step. Find the twentieth term and the sum of the first twenty.

The twentieth term is 7+19×4=83. The sum is 20×(7+83)/2=20×45=900. Check the first few by hand: 7+11+15+19=52, and the formula on four terms gives 4×(7+19)/2=52.

Now you. A sequence starts at 5 and rises by 3. Find the thirtieth term and the sum of the first thirty.

Answer

a30=5+29×3=92, and S30=30×(5+92)/2=15×97=1455.

Geometric sequences and the shift trick

A sequence is geometric when consecutive terms have a constant ratio r. Each step multiplies, so an=a1rn-1, and the exponential functions of the earlier lesson are the continuous version of the same idea.

Summing needs a different device, and it is worth seeing because the same move settles the infinite case. Write

Sn=a+ar+ar2+dots+arn-1

Multiply the whole thing by r, which shifts every term along one place: rSn=ar+ar2+dots+arn. Subtract the second from the first and everything in the middle cancels, leaving only the first term of one and the last of the other: Sn-rSn=a-arn. Factor and divide, which is legal provided r1:

Sn=a1-rn1-r

Example. Sum the first ten terms of 3,6,12,24,dots

Here a=3 and r=2, so S10=3(1-210)/(1-2)=3(1024-1)=3069, and the tenth term itself is 3×29=1536. Note that the last term is nearly half the whole sum, which is characteristic: in a geometric series with r=2 every term exceeds the total of all the terms before it.

Now you. Sum the first eight terms of 5,15,45,dots

Answer

a=5, r=3, so S8=5(38-1)/(3-1)=5×6560/2=16{,}400, and the eighth term is 5×37=10{,}935.

Adding infinitely many numbers

Now let n grow without bound in that formula. The only part that depends on n is rn, and its behaviour was settled two lessons ago: if lvertrrvert<1 then repeated multiplication by r drives the term toward zero. So the sum settles down to a definite value,

S=a1-rprovided lvertrrvert<1

and if lvertrrvert1 the terms do not shrink and the partial sums run away without limit. So 6+2.4+0.96+dots, with r=0.4, totals 6/0.6=10, and no matter how many terms are taken the running total stays below 10 and closes on it.

This retires a loose end from the first lesson. The repeating decimal 0.272727dots is the series 27/100+27/10000+dots, geometric with a=27/100 and r=1/100, so its total is (27/100)/(99/100)=27/99=3/11, exactly the answer the shifting trick gave there. The same argument applied to 0.999dots gives (9/10)/(9/10)=1. That is not an approximation and it is not a paradox: the notation 0.999dots means the value the partial sums close on, and that value is 1.

The requirement lvertrrvert<1 is essential and its failure is not subtle. Zeno's paradox of Achilles and the tortoise is a geometric series with a ratio below one, so the infinitely many gaps are crossed in a finite time and there is no paradox at all. A series with r=1.5 has partial sums that grow forever, and writing down a formula's answer for it produces a number that means nothing.

What a loan costs

Compound interest and geometric series together price every fixed-payment loan, and the derivation is short enough to do rather than quote. Borrow P at a monthly rate i and repay a fixed amount M at the end of each of n months. Each payment must cancel the debt it is settling, and a payment made k months from now is worth M(1+i)-k today, since that sum invested now would grow to M by then. Setting the total present value of the payments equal to the loan gives

P=Mbig[(1+i)-1+(1+i)-2+dots+(1+i)-nbig]

The bracket is geometric with first term and ratio both (1+i)-1. Summing it and tidying gives the standard formula

M=Pi1-(1+i)-n

Example. A mortgage of 200{,}000 pounds at a nominal 4.2 per cent a year runs for 25 years with monthly payments. What is the payment, and what does the loan cost in total?

The monthly rate is i=0.042/12=0.0035 and n=300. Then (1.0035)-300=0.350580, so the denominator is 0.649420 and M=200000×0.0035/0.649420=1077.88 pounds a month. Over 300 months that is 323{,}365 pounds, so the interest is 123{,}365 pounds, more than sixty per cent of the sum borrowed.

Now you. Find the monthly payment and the total interest on 150{,}000 pounds at 5.4 per cent over 20 years.

Answer

i=0.0045 and n=240, giving M=1023.38 pounds a month. The total paid is 245{,}611 pounds, so the interest is 95{,}611 pounds.

The same formula run the other way prices saving. Paying 200 a month into an account at 4 per cent for thirty years contributes 72{,}000 of capital and accumulates to 138{,}810, the difference being interest on interest.

When shrinking terms are not enough

It is tempting to conclude that a series converges whenever its terms shrink toward zero. It does not, and the standard counterexample is the harmonic series 1+1/2+1/3+1/4+dots, whose terms clearly shrink to nothing.

Nicole Oresme showed around 1350 that it grows without limit, by an argument that still works. Group the terms as 1/3+1/4, then 1/5 through 1/8, then 1/9 through 1/16, doubling the block each time. Every term in a block is at least as large as the last one in it, and there are enough of them that each block totals at least 1/2. Infinitely many blocks each contributing at least a half cannot add to anything finite.

The growth is extraordinarily slow, which is why it fools people. The first thousand terms total 7.4855 and the first million total 14.3927; passing 10 takes 12{,}367 terms. Slow growth is still growth, and the series has no sum. Deciding which series converge, in general, is one of the main tasks of calculus, and this lesson has settled only the geometric case.

Sequences that define themselves

Some recursive sequences have no simple closed form and are still worth studying. Fibonacci's, from the Liber Abaci of 1202, is F1=F2=1 with Fn+1=Fn+Fn-1, giving 1,1,2,3,5,8,13,21,dots

Its terms grow, and the ratio of consecutive terms settles: F22/F21=17711/10946=1.618034, against the golden ratio (1+5)/2=1.618034. That is not a coincidence and it can be proved, as can a closed form for Fn involving powers of the golden ratio, though both need tools beyond this lesson.

What the computation establishes is weaker than it looks, and this is the point on which the course now turns. Twenty terms have been checked. Nothing checked so far rules out the ratio drifting away at the two hundredth term, or the arithmetic sum formula failing at some enormous n, or a pattern that holds for every case anyone has patience to test and fails after that.

Such failures are not hypothetical. The polynomial n2+n+41, noted by Euler, produces a prime for every n from 0 to 39 and then fails at n=40, where it gives 1681=412. Forty successes is a great deal of evidence and no proof at all.

So the closed forms of this lesson rest on arguments, the pairing and the shift, that were performed on a written-out list with dots in the middle. Those arguments are convincing, and being convincing is not the standard. What the standard is, and how to meet it for a claim about every natural number at once, is the last lesson.

What a proof is

The previous lesson derived a sum formula by writing a list forwards and backwards with dots in the middle, and the first lesson asserted that no fraction squares to 2 without giving the argument. Both are believed by every mathematician, and neither has yet been established here.

This lesson is about the difference between believing and establishing. It assumes everything the course has covered, and it settles the two open claims using techniques that are the whole toolkit: direct argument, contrapositive, contradiction and induction.

Why evidence is not enough

A mathematical claim usually quantifies over infinitely many cases. "The sum of the first n odd numbers is n2" is not one statement but infinitely many, one for each n, and checking a hundred of them leaves infinitely many unchecked.

That would be pedantry if patterns held reliably. They do not. Euler's n2+n+41 from the previous lesson gives a prime for every n from 0 to 39 and fails at 40. Fermat conjectured in 1640 that every number of the form 22n+1 is prime, having checked n up to 4, and Euler showed in 1732 that the next one factors: 232+1=4{,}294{,}967{,}297=641×6{,}700{,}417. Five successes and then failure.

A proof is an argument that the claim cannot fail, given the assumptions, for any case at all. It replaces sampling with necessity. This is the trade that makes mathematics different from the empirical sciences: the conclusions are narrower, since they hold only under stated assumptions, and within that scope they are permanent. A theorem of Euclid's is as true now as it was in 300 BCE, which cannot be said of any measurement of that era.

Statements, and the asymmetry of the counterexample

Two shapes of claim behave differently under testing. A claim about all cases, "every prime above 2 is odd", cannot be established by examples but can be destroyed by a single one that fails. A claim about some case, "there exists a prime between 90 and 100", is established by one example, namely 97, and refuting it requires an exhaustive argument.

So the first thing to do with a suspicious claim about all cases is to hunt for a counterexample, and finding one ends the matter with no further work.

Example. Is it true that 2p-1 is prime whenever p is prime?

Test the small cases: p=2 gives 3, p=3 gives 7, p=5 gives 31, p=7 gives 127, all prime, and the pattern looks solid. At p=11 it gives 2047, which is 23×89. The claim is false, and one line of arithmetic settles it. These numbers are the Mersenne primes when they are prime, and which primes p produce them is unsolved to this day.

Now you. Is it true that n2-n+11 is prime for every natural number n? Either find a counterexample or say why the search should start where it does.

Answer

It fails at n=11, where the expression is 121-11+11=121=112. The search should start there because every term is a multiple of 11 when n is, which is the same mechanism that breaks Euler's polynomial at n=41.

Direct proof

The plainest structure assumes the hypothesis and reasons forward to the conclusion, using definitions and established results. The essential move is to replace a word with its definition, since a definition is the only thing that lets an argument get started.

Claim: the sum of two even numbers is even. An even number is by definition 2k for some integer k. So take two of them, 2k and 2m. Their sum is 2k+2m=2(k+m) by distributivity, and k+m is an integer, so the sum has the required form and is even. The proof is three lines and it covers infinitely many pairs at once, which is exactly what checking cases could not do.

The same shape proves that the product of two odd numbers is odd. Odd means 2k+1, so the product is (2k+1)(2m+1)=4km+2k+2m+1=2(2km+k+m)+1, which is odd. Notice that the algebra doing the work is the distributivity of the second lesson, and that the derivation of (-1)(-1)=1 given there was itself a direct proof.

Contrapositive, and the converse that is not the same

Some claims resist a direct attack. Show that if n2 is even then n is even, and the hypothesis n2=2k gives nothing useful to factor.

Every implication "if P then Q" is logically identical to its contrapositive, "if not Q then not P". If rain implies wet ground, then dry ground implies no rain, and the two say the same thing. So instead prove: if n is odd then n2 is odd. That is direct and immediate, since n=2k+1 gives n2=4k2+4k+1=2(2k2+2k)+1, which is odd. The original claim follows with no further work.

The converse, "if Q then P", is a different statement and is not implied. Wet ground does not imply rain. Confusing an implication with its converse is the most frequent error in reasoning inside mathematics and outside it, and the discipline is to state which of the two is being claimed. When both hold, the statement is written "if and only if", and it requires two proofs. The factor theorem in the polynomial lesson was such a claim, and it needed both directions.

Contradiction, and the diagonal of the square

Proof by contradiction assumes the claim is false and derives an impossibility, so the assumption cannot stand. It is the technique the first lesson promised, and here is the debt paid.

Claim: there is no rational number whose square is 2.

Suppose there were. Then 2=p/q for integers p and q with no common factor, since any common factor can be cancelled first, and this reduction is where the contradiction will eventually be trapped. Squaring gives 2=p2/q2, so

p2=2q2

The right side is even, so p2 is even, so by the contrapositive result proved above p is even. Write p=2m. Substituting gives 4m2=2q2, so q2=2m2, and the same argument makes q even too.

But p and q were assumed to have no common factor, and both are now even, so both are divisible by 2. That is a contradiction, so the assumption that 2 is rational is false. The diagonal of a unit square has a length no fraction names, which is what the Pythagoreans found and what forced the real line into existence in the first lesson.

Euclid's proof that the primes never run out has the same shape and is worth knowing. Suppose there were finitely many, p1 through pk. Form N=p1p2pk+1. Dividing N by any prime on the list leaves remainder 1, so no listed prime divides it. Yet every integer above 1 has a prime factor, so N has one, and it is not on the list, contradicting the assumption that the list was complete. Note what the argument does not claim: N need not itself be prime. With the first six primes, N=30031=59×509, and both factors are missing from the list, which is all the proof requires.

Induction, and the two open sum formulas

Contradiction and direct argument still do not reach a claim indexed by every natural number. For that there is mathematical induction, and the picture is a line of dominoes: knock over the first, and guarantee that each one knocks over its neighbour, and all of them fall, however many there are.

Formally, to prove a statement S(n) for every natural n, prove two things. The base case: S(1) is true. The inductive step: if S(n) is true then S(n+1) is true. Both are finite tasks, and together they establish infinitely many statements. The step is not the assumption that S(n) is true for all n, which would be circular; it is the proof of a conditional, and that distinction is the whole of the technique.

Example. Prove that the sum of the first n odd numbers is n2.

Base case: for n=1 the sum is 1 and 12=1, so it holds. Inductive step: assume 1+3+dots+(2n-1)=n2 for some particular n. The next odd number is 2n+1, so the sum to n+1 terms is n2+(2n+1) by the assumption, and n2+2n+1=(n+1)2, which is the claim for n+1. Both parts hold, so the formula is true for every n. The algebra in the step is the perfect square from the third lesson, doing the essential work.

Now you. Prove by induction that 1+2+dots+n=n(n+1)/2, the formula the previous lesson obtained by pairing.

Answer

Base case: at n=1 the sum is 1 and 1×2/2=1. Step: assume the sum to n is n(n+1)/2. Then the sum to n+1 is n(n+1)/2+(n+1)=(n+1)big(n/2+1big)=(n+1)(n+2)/2, which is the formula with n+1 in place of n. So it holds for every n.

Induction also handles the geometric sum of the previous lesson, whose derivation by shifting and subtracting relied on dots standing in for unwritten terms. Assume Sn=a(1-rn)/(1-r). Adding the next term arn gives a(1-rn)/(1-r)+arn, and putting it over the common denominator gives a(1-rn+rn-rn+1)/(1-r)=a(1-rn+1)/(1-r), which is the formula one step on. With the base case S1=a checked directly, the formula is proved for every n, and every loan calculation in the previous lesson now stands on something firmer than a persuasive picture.

Example. Prove that 2n>n2 for every n5.

Base case: at n=5, 32>25. Note that the claim is false at n=4, where both are 16, which is why the base is placed at 5. Step: assume 2n>n2 for some n5. Then 2n+1=2×2n>2n2, so it is enough to show 2n2(n+1)2, that is n2-2n-10. By the quadratic formula that holds for n1+2=2.414, and n5 comfortably satisfies it. So the claim propagates.

Now you. Prove that n3-n is divisible by 3 for every natural n.

Answer

Base case: at n=1 the value is 0, which is divisible by 3. Step: assume n3-n=3k. Then (n+1)3-(n+1)=n3+3n2+3n+1-n-1=(n3-n)+3(n2+n)=3k+3(n2+n)=3(k+n2+n), a multiple of 3.

What proof does not give you

Three honest limits. A proof establishes a conclusion from assumptions, and it says nothing about whether the assumptions describe anything real. Euclidean geometry is proved from Euclid's postulates, and physical space does not obey them.

A proof can be wrong, and long ones sometimes are. Errors surface in refereeing, and famous arguments have collapsed and been repaired: Wiles' first version of the proof of Fermat's last theorem, announced in 1993, had a gap that took a further year and a collaborator to close.

And no system of axioms strong enough to describe arithmetic can prove every true statement about the natural numbers. Gödel established that in 1931, and it is not a loophole any working mathematician trips over, but it does mean that "provable" and "true" are not the same word.

Where this ends and the next thing starts

The subject is now closed. Numbers were built from counting to the continuum, algebra was reduced to a short list of laws, equations were solved, functions were defined, drawn and inverted, and polynomials, exponentials, logarithms and the circular functions were each derived and put to work on data. The two arguments the course deferred have been given, and the technique that gives them is the one that would prove the rest.

What remains untouched is the question every one of these functions raises and none of them answers: how fast is it changing right now, and what does the area under it total. Both are limits, and both require the completeness of the real line established in the first lesson. That is calculus, and everything needed to start it is now in place.

Mathematical Foundations, from libre.university