Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Music Theory

Read a score and say what it is doing: rhythm, intervals, scales and keys, chords, and why a progression pulls the way it does.

Notes and pitch

A note is air pressure rising and falling at a steady rate, and everything in this course is eventually built on that one physical fact. This lesson turns it into numbers, and the numbers into the twelve named pitches the rest of the subject uses.

A note is a rate

Pluck a guitar string and it swings back and forth some hundreds of times a second, pushing the air with it. The number of complete swings per second is the frequency, measured in hertz. Frequency is what the ear reports as pitch: raise it and the note sounds higher, lower it and the note sounds lower, and nothing else about the sound changes that judgement much.

The reference point is a convention rather than a discovery. Since a conference in London in 1939 the standard has been that the A above middle C sits at 440 Hz, written A4 = 440 Hz. It is not sacred. Many European orchestras tune to 442 Hz, which is 7.9 cents above the standard by the measure defined later in this lesson, and ensembles playing baroque music commonly tune to 415 Hz, which is 101.3 cents below it, almost exactly one semitone. A piece is not transposed by moving the standard; everything moves together, and the relations between the notes survive untouched.

That is the first real clue about how music works. What a listener recognises is not a set of absolute frequencies but a set of ratios between them. Play a tune starting on 440 Hz, then play it starting on 330 Hz, and almost everyone hears the same tune. Play it with every frequency shifted up by 50 Hz instead and it is unrecognisable, because adding a constant destroys the ratios while multiplying by one preserves them.

The octave is exactly a doubling

One ratio stands apart. Sound 440 Hz and 880 Hz together and they do not sound like two notes so much as one note in two places. Every musical tradition that has been studied treats the pair as a kind of equivalence, and Western notation goes so far as to give them the same letter: both are A. The interval between them is the octave, and it is the ratio 2:1 to whatever precision you care to measure.

Why that ratio and no other is a fact about how a physical body vibrates. A string fixed at both ends can vibrate as a whole, but it can also vibrate in two halves at twice the rate, in three thirds at three times the rate, and so on, and in practice it does all of these at once. The sound arriving at your ear is a stack of frequencies f, 2f, 3f, 4f, ... called the harmonic series, and their relative strengths are most of what makes a violin sound different from a clarinet playing the same note.

Take the low C on a cello, C2 at 65.41 Hz, and list the first eight members of its stack: 65.41, 130.81, 196.22, 261.63, 327.03, 392.44, 457.84, 523.25 Hz. Now sound the note an octave above, C3 at 130.81 Hz, whose own series is 130.81, 261.63, 392.44, 523.25 and so on. Every frequency in the upper note's series is already present in the lower note's. Nothing new arrives, which is exactly what "the same note again" means physically. That the ear treats the octave as an identity is not mysticism; it is the ear noticing that one spectrum is a subset of another.

Example. A tuning fork sounds at 256 Hz. What are the frequencies two octaves above and one octave below?

Each octave is a factor of 2, so two octaves up is 256×2×2=1024 Hz and one octave down is 256/2=128 Hz. Octaves multiply. They never add: the interval from 256 to 512 Hz spans 256 Hz, and the identical-sounding interval from 512 to 1024 Hz spans 512 Hz.

Now you. A double bass string sounds at 41.20 Hz. Give the frequency three octaves above it, and say how many hertz wide that whole span is.

Answer

Three octaves up is 41.20×23=41.20×8=329.6 Hz. The span is 329.6-41.2=288.4 Hz, but that figure is musically useless: the same three octaves starting from 82.40 Hz would span 576.8 Hz and sound identical.

The other consonances are the other small ratios

If the octave is 2:1, the next candidates are 3:2, 4:3, 5:4 and 6:5, and they are exactly the intervals Western music treats as consonant. The connection runs through the harmonic series again. Two notes whose frequencies are in a small whole-number ratio share many upper partials, so the combined spectrum stays tidy; two notes in a ratio like 45:32 share almost nothing below the forty-fifth partial, and the result is the interval people call harsh.

You can read these ratios straight off the C2 series above. The third partial, 196.22 Hz, sits at 3:2 above the second partial at 130.81 Hz, and that pair is the interval later called a perfect fifth. The fifth partial at 327.03 Hz sits at 5:4 above the fourth at 261.63 Hz, a major third. The sixth over the fifth, 392.44 over 327.03, is 6:5, a minor third. Stacking a major third and a minor third gives 54×65=32, which is why a major triad, the ordinary chord that arrives later in this course, is built the way it is. The chord was in the physics before anybody wrote it down.

The seventh partial is a warning. At 457.84 Hz it lies close to the note keyboards call B flat, but a keyboard B flat is 466.16 Hz, and the two differ by 31.2 cents, roughly a third of a semitone. It is genuinely a different note, and no amount of tuning a piano will produce it. Western notation has no symbol for it, which is one of the first honest limits of the system this course teaches.

Twelve steps, and why it is twelve

So far there are ratios but no scale. The historical route to one is to take the strongest interval after the octave, the fifth at 3:2, and pile it up. Start at some note and go up a fifth twelve times, and the total ratio is

(32)12=129.746

Going up seven octaves instead gives 27=128. Those two numbers are not equal, but they are close: twelve fifths overshoot seven octaves by a factor of 129.746/128=1.013643, an excess known as the Pythagorean comma. It is about a quarter of a semitone, small enough that the cycle very nearly closes and large enough to be plainly audible as out of tune.

That near-miss is the whole argument for twelve. Because twelve fifths land almost exactly seven octaves higher, a chain of fifths run through twelve steps produces twelve distinct notes per octave and then starts repeating, to within a comma. Try other chain lengths and the arithmetic is worse: five fifths or seven fifths give coarser divisions, and the next genuinely good closure is not until 41 or 53 steps, which no keyboard player would thank you for. Twelve is the smallest number of equal parts that keeps a good fifth, and Western music took it.

Equal temperament

The comma still has to go somewhere. Tuning by pure ratios, called just intonation, gives one gorgeous key and progressively worse ones as you move away from it, because the comma piles up in whichever intervals were left over. The eighteenth-century solution was to abolish the problem by fiat: make all twelve steps identical, and let every interval except the octave be slightly wrong.

If twelve equal steps multiply to an octave, each step is a ratio r with r12=2, so

r=21/12=1.059463

That step is the semitone, and this tuning is equal temperament. Everything follows from it. The note n semitones above A4 has frequency

f=440×2n/12

with n negative for notes below. Middle C is nine semitones below A4, so f=440×2-9/12=261.63 Hz. The A at the bottom of a piano is 48 semitones below, giving exactly 27.5 Hz, and the top C is 39 semitones above, giving 4186 Hz.

Example. What is the frequency of the E a perfect fifth above A4, and how far is it from the pure 3:2?

A fifth is seven semitones, so f=440×27/12=659.26 Hz. A pure fifth above 440 Hz would be 440×1.5=660 Hz. The tempered fifth is flat by 0.74 Hz, about one part in 900.

Now you. Find the frequency of the note four semitones above middle C, which is E4, and compare it with the pure 5:4 major third above middle C at 261.63 Hz.

Answer

f=261.63×24/12=329.63 Hz. The pure third would be 261.63×1.25=327.03 Hz, so the tempered third is sharp by 2.60 Hz, a much bigger error than the fifth's.

Cents, the unit that makes errors comparable

Comparing 0.74 Hz at one pitch with 2.60 Hz at another is meaningless, because a fixed number of hertz is a large interval down low and a negligible one up high. Intervals are ratios, so the honest unit is logarithmic. The cent divides the equal-tempered semitone into 100 parts, so an octave is 1200 cents, and the size in cents of a frequency ratio R is

n=1200log2R

Now the compromise can be audited. A pure fifth is 1200log2(3/2)=701.955 cents, while equal temperament offers 700, so the tempered fifth is 1.96 cents flat: inaudible in practice. A pure major third is 1200log2(5/4)=386.31 cents against a tempered 400, so the tempered third is 13.69 cents sharp, which is about a seventh of a semitone and very much audible. The minor third is worse in the other direction, 315.64 cents pure against 300 tempered, 15.64 cents flat.

This is the price list, and it is worth knowing rather than glossing. Equal temperament buys the ability to play in every key on one instrument, and pays for it with thirds that are noticeably sharp. String quartets and unaccompanied choirs, which are free to adjust, routinely do not play equal-tempered thirds at all.

Hearing the error

The size of the error is checkable by ear, not just by arithmetic, because two nearly-equal frequencies produce beats: the loudness pulses at a rate equal to the difference between them.

Sound middle C at 261.63 Hz and the tempered E4 at 329.63 Hz together. The fifth partial of the C lies at 5×261.63=1308.15 Hz and the fourth partial of the E lies at 4×329.63=1318.52 Hz. Those two nearly coincide, and their difference is 10.37 Hz, so the chord throbs about ten times a second. Do the same with the fifth C4 and G4: the third partial of C is 784.89 Hz, the second partial of G is 783.99 Hz, and the beat rate is 0.89 Hz, one slow pulse per second. That contrast, ten beats a second against one, is the 13.69 cents against 1.96 cents made audible, and it is why piano tuners tune fifths by counting slow beats and check thirds by counting fast ones.

Example. A piano's A4 is tuned to 440.0 Hz and its A3 to 219.5 Hz. What do you hear?

The octave should be exactly 2:1, so A3 should be 220.0 Hz. Its second partial is 2×219.5=439.0 Hz against the A4 at 440.0 Hz, a difference of 1.0 Hz, so the octave beats once per second. Tuning by ear means adjusting until that beating stops.

Now you. Two flutes play a note nominally at 587.33 Hz, but one is 1.5 Hz sharp. How many beats per second, and how many cents apart are they?

Answer

The beat rate is just the difference, 1.5 beats per second. In cents, 1200log2(588.83/587.33)=4.4 cents, which is small on the page and impossible to ignore in the room.

What the numbers do not decide

Three limits are worth stating before moving on. Equal temperament is a choice, not a law of hearing: 19, 31 and 53 equal divisions of the octave all have their advocates, and 31 gives a major third within 0.8 cents of pure. Much of the world's music does not use twelve at all, and maqam and dastgah traditions use intervals that fall between the keys of a piano rather than on them. And frequency is not the whole of a sound: two instruments at 440 Hz differ in the relative strength of their partials and in how the note begins, which is timbre, and this course says nothing about it.

What has been established is that a pitch is a number and that the useful relations between pitches are ratios. Nobody reads numbers off a page, though. The next lesson turns the printed dot into the named frequency.

Reading the staff

The previous lesson ended with a pitch as a number, and no reader has ever opened a score and found numbers in it. What a score contains is dots on lines, and this lesson is the decoder.

Height on the page means pitch

Western notation makes one very good decision at the start: it maps pitch onto vertical position, so higher on the page is higher in frequency. That single choice is why a melodic shape is visible at a glance and why a rising line looks like a rising line.

The grid it maps onto is the staff, five horizontal lines with four spaces between them. A notehead sits either centred on a line or centred in a space, and those nine positions, plus the ones extended above and below, are the only places a note can go. Moving up one position, from a line to the space above it, is one step through the letter names: A to B, B to C, and so on round the alphabet, which runs A B C D E F G and then starts again.

That last point matters more than it looks. The staff counts letters, not semitones. Adjacent positions on the staff are sometimes a tone apart and sometimes a semitone apart, because the seven letters are unevenly spread across the twelve semitones of the octave. From B to C and from E to F is one semitone; every other pair of adjacent letters is two. The staff is a diagram of the seven-letter system, and the missing five pitches have to be bolted on with accidentals later in this lesson. Why the letters are spaced that way is not arguable from the staff alone: it is a fact about the major scale, which the lesson on scales takes up.

Clefs fix which pitch

Five lines with no further information say nothing absolute. They give you the shape of a tune but not its pitch, because nothing declares what the bottom line is. That declaration is the clef, printed at the head of every staff.

The treble clef is a stylised letter G, and the point of its inner spiral is the whole content of the symbol: the line it curls around is G4, the G above middle C. Everything else follows by counting positions. From that G, the line below is E4 and the line above is B4, so the five lines of a treble staff, bottom to top, are E4, G4, B4, D5, F5, and the four spaces between them are F4, A4, C5, E5.

The bass clef is a stylised letter F, and its two dots straddle the line that is F3, the F below middle C. Counting from there gives lines G2, B2, D3, F3, A3 and spaces A2, C3, E3, G3.

Both exist because one staff is not enough. Music in common use spans roughly seven octaves, and a single staff covering it would need dozens of lines that no eye could count. Five is close to the limit for reading a position without counting. Two clefs, placed a good distance apart, cover the whole practical range with staves that stay readable, and each singer or instrument uses the one that keeps its notes on or near the lines. There are also two moveable C clefs, which mark middle C at whichever line they are centred on: the alto clef, with middle C on the middle line, is what viola players read, and the tenor clef, with middle C on the fourth line, appears in the upper register of cello, bassoon and trombone parts. They exist for the same reason, which is to avoid ledger lines.

The grand staff and middle C

Keyboard music uses both clefs at once, braced together as a grand staff, with the right hand usually reading the treble and the left the bass. The gap between them is not arbitrary. The top line of the bass staff is A3 and the bottom line of the treble staff is E4, and continuing the pattern upward through the gap gives a space at B3, a line at C4 and a space at D4. That middle line is the one the printer leaves out, and the note it would carry is C4, middle C.

That is why middle C is drawn on a short line of its own, one ledger line below the treble staff or one above the bass staff, and why both drawings mean the same key on the piano. Think of the grand staff as an eleven-line system with the middle line rubbed out and restored one note at a time, and the whole layout stops needing to be memorised.

Example. A note sits in the second space from the bottom of the treble staff. Name it, and give its frequency.

The treble spaces upward are F4, A4, C5, E5, so the second is A4. From the previous lesson, A4 is the tuning reference at 440 Hz exactly.

Now you. A note sits on the top line of the bass staff. Name it and give its frequency.

Answer

The bass lines upward are G2, B2, D3, F3, A3, so the top line is A3. It is twelve semitones below A4, so 440×2-12/12=220 Hz.

Naming which octave

"A" names seven different pitches on a piano, so a written note needs an octave label as well as a letter. The standard is scientific pitch notation, which appends a number: A4 is the tuning A, A3 the octave below, A5 the octave above.

The one trap is where the number changes. It does not change at A, despite the alphabet starting there; it changes at C. The sequence running upward is B3, C4, D4 ... B4, C5, so middle C begins octave 4 and the A four notes above it is A4. A note written B3 is one semitone below C4 and shares no digit with it.

The same numbering underlies the note numbers used by MIDI, where middle C is 60 and A4 is 69, one number per semitone. That gives a direct route from a printed dot to a frequency: read the letter and octave off the staff, convert to a note number m, and apply

f=440×2(m-69)/12

which is the formula from the previous lesson with its reference shifted. Be aware that software disagrees about the names, though never about the numbers: some manufacturers label note 60 as C3 rather than C4, so a synthesiser's display and a theory textbook can be an octave apart in words while agreeing exactly in pitch.

Ledger lines and octave signs

Notes beyond the staff get ledger lines, short line segments continuing the pattern above or below. They are counted, and counting is slow, which is why more than three or four in a row is considered bad engraving.

Example. How many ledger lines does C6 need above a treble staff?

The top line is F5. Continuing upward, the space above it is G5, then the first ledger line is A5, the space above that is B5, and the second ledger line is C6. So two ledger lines, and the notehead sits on the upper one.

Now you. How many ledger lines does C2 need below a bass staff, and what note sits on the first of them?

Answer

The bottom line is G2. Below it, the space is F2, the first ledger line is E2, the space below that is D2, and the second ledger line is C2. Two ledger lines, with E2 on the first.

The alternative is an octave sign. 8va written above a passage means play it an octave higher than written, 8vb below a passage means an octave lower, and 15ma means two octaves higher, the 15 counting inclusively in the way intervals are counted in a later lesson. A dashed line shows how far the instruction runs, and loco cancels it. Piccolo, glockenspiel and double bass parts would be unreadable without them.

Accidentals, and how long they last

Seven letters cannot name twelve pitches, so five of the twelve are written as a letter plus a sign. There are five signs. The sharp raises a note one semitone, the flat lowers it one semitone, the natural cancels either, the double sharp raises by two semitones, and the double flat lowers by two. In printed music they are the symbols ♯, ♭, ♮ and their doubles; in running text, including everywhere in this course and in every quiz answer, they are written as they are typed, so F sharp is F# and B flat is Bb.

An accidental is printed to the left of the notehead, but it applies forward, and its life is precisely defined: it lasts until the end of the bar, and it applies only to that letter in that octave. Write F# on the first beat of a bar, and every F on that same line or space for the rest of the bar is sharp without being marked again. The barline cancels it. An F an octave higher in the same bar is strictly unaffected, though modern editors nearly always add a cautionary accidental in brackets rather than trust the reader, and you should read the bracket as politeness, not as new information. A note tied across a barline keeps its accidental for as long as the tie lasts, since it is one sounding note rather than two.

Example. A bar in the key of C major contains, in order, F, F#, F, G, and then a barline followed by F. Which of those F notes sound as F#?

The first F is plain, because nothing has altered it. The F# is sharp by its own sign. The third F is also sharp, since the accidental holds until the barline. The F after the barline is plain again.

Now you. A bar contains Bb on beat 1, then B natural marked on beat 2, then B on beat 4, and the following bar opens with B. Which of them sound as B natural?

Answer

Beats 2 and 4 of the first bar. The natural sign cancels the flat for the rest of that bar, so the B on beat 4 is natural too. The B in the next bar reverts to whatever the key signature says, which in a key with a Bb in its signature means B flat again, a point the lesson on keys will make properly.

Two names for one sound

Because a sharp raises and a flat lowers, more than one spelling can land on the same pitch. C# and Db are the same key on a piano and the same frequency in equal temperament, as are F# and Gb, and less obviously E# and F, or Cb and B. Such pairs are enharmonic equivalents.

It is tempting to conclude that the spelling is decoration. It is not, for two reasons. The practical one is that the staff position differs: C# sits on the C line with a sharp, Db sits on the D line with a flat, and the choice changes what the player's hand does and how the passage reads. The musical one is that spelling records function, and function is what the second half of this course is about. In the key of D major the note between C and D is C#, the leading note that pulls up to the tonic, and writing it Db would say the opposite thing about where the music is going.

There is one place the equivalence is not even exactly true. In equal temperament C# and Db are identical by construction, since the tuning was flattened until they were. In just intonation or in the unequal temperaments used before it, they are different frequencies, and violinists playing without a keyboard still bend them apart, playing a leading note high and a flattened note low.

What the staff does not say

A page of noteheads on lines now decodes to a list of exact frequencies. What it does not decode to is music, because nothing so far says when any of them happens or for how long. That is not a small omission: rhythm is the half of notation this lesson has entirely ignored, and the next one takes it up, starting from a single note and cutting it in half.

Note values and duration

A note on a staff says which frequency and nothing about when, and a page of pitches with no durations is not yet music. This lesson supplies the other half of a notehead's meaning.

Duration is a fraction, not a time

The first thing to accept is that written durations are relative. A symbol on the page says how long a note lasts compared with the other notes around it, and says nothing about seconds until a tempo mark is added at the end of this lesson. That separation is deliberate and useful: it lets the same rhythm be performed fast or slow without rewriting a single notehead.

The unit at the top of the system is the whole note, and every other value is reached by halving it. Halve a whole note and you have a half note, halve again for a quarter note, again for an eighth, a sixteenth, a thirty-second, a sixty-fourth. British names describe the same objects: semibreve, minim, crotchet, quaver, semiquaver, demisemiquaver, hemidemisemiquaver. Both naming systems appear in print, and the American one is used here because it says the arithmetic out loud.

The symbols are built to make the halving visible. A whole note is a hollow notehead with no stem. Add a stem and it is a half note. Fill the head in and it is a quarter note. From there, each further halving adds a flag to the stem: one flag for an eighth, two for a sixteenth, three for a thirty-second. So a filled head with two flags is a sixteenth note, and you can read the value off the drawing without recognising the shape as a whole. Above the whole note sits one survivor of older notation, the double whole note or breve, worth two whole notes and now rare outside choral music and final chords.

Rests are the same tree, silent

Every note value has a matching rest, drawn differently but worth exactly the same. A whole rest is a small filled rectangle hanging below the fourth line of the staff, a half rest is the same rectangle sitting on the third line, a quarter rest is a zigzag, and eighth and shorter rests carry flags in the same pattern the notes do: one flag for an eighth rest, two for a sixteenth.

The two rectangles are worth learning properly, because hanging and sitting are the whole difference between two beats and four. There is also one deliberate inconsistency: a whole rest, in most contexts, means a whole bar of silence rather than four quarters of it, so in a bar of three quarters it is worth three, not four. That is a convention about bars, which the next lesson defines, and it is the one place where the halving tree is quietly overruled by context.

Dots, and what halving cannot reach

Halving alone can only produce durations that are powers of two: one whole, one half, one quarter. It can never produce a duration of three quarters, because 3 is not a power of 2 times anything in the tree. Since music in three is at least as old as music in two, notation needs a repair, and the repair is the dot.

A dot to the right of a notehead adds half of that note's value to it. A dotted half note is worth 2+1=3 quarter notes. A dotted quarter is worth one and a half quarters, which is three eighths. A dotted whole note is worth six quarters.

A second dot adds half of the first dot, that is, a quarter of the original. A double-dotted half note is therefore worth 2+1+0.5=3.5 quarter notes. The pattern is a geometric series: with n dots the total is

1+12+14++12n=2-12n

times the undotted value. One dot gives 2-12=1.5 times, two dots 1.75 times, three dots 1.875 times. Notice what the formula says about the limit: no number of dots ever reaches twice the value, which is why a dotted note is always shorter than the next note up the tree and never equal to it. Three dots is the practical maximum and is already unusual.

Example. How many quarter notes is a double-dotted whole note worth?

A whole note is 4 quarters, and two dots multiply by 2-14=1.75, so 4×1.75=7 quarter notes. The same result by addition: 4+2+1=7.

Now you. How many eighth notes is a double-dotted quarter note worth?

Answer

A quarter note is 2 eighths, and 2×1.75=3.5 eighths. By addition, 2+1+0.5=3.5, so it equals seven sixteenth notes.

Ties, and durations no symbol can spell

Dots close many of the gaps but not all of them. A duration of five sixteenths, for instance, is out of reach: a quarter note is four sixteenths, a dotted quarter is six, and nothing in between exists. The general repair is the tie, a curved line joining two noteheads of the same pitch, which instructs the player to sound the first and hold it through the second without rearticulating. Five sixteenths is a quarter note tied to a sixteenth.

Because any duration can be written as a sum of tree values, ties plus dots can express anything. They are also the only way to write a note that continues across a barline, since a notehead cannot be printed in two bars at once, and that is their commonest use by far.

A tie looks exactly like a slur, which is also a curved line, and means something entirely different: a slur joins notes of different pitches and asks for them to be played smoothly and connectedly, each still sounded. The pitches tell you which is which. Two identical noteheads joined by a curve is a tie and one sound; a curve over a rising phrase is a slur and as many sounds as there are notes.

Example. Write a duration of seven sixteenth notes.

Check the dots first. A quarter is 4 sixteenths, so a double-dotted quarter is 4×1.75=7 sixteenths exactly. No tie is needed at all, which is worth checking for before reaching for one.

Now you. Write a duration of nine sixteenth notes.

Answer

Dots cannot do it: a half note is 8 sixteenths, a dotted half is 12, and a double-dotted half is 14. So tie a half note to a sixteenth, 8+1=9. A dotted quarter tied to a dotted eighth also gives 6+3=9, and both are correct arithmetic, though the first is far easier to read.

A tempo mark makes it seconds

Now the fractions become time. A metronome mark at the head of a piece names a note value and a number of them per minute: quarter note = 120 means 120 quarter notes pass in a minute, so each quarter note lasts 60/120=0.5 s. In general, a note worth v quarter notes at a mark of B quarter notes per minute lasts

t=v×60Bseconds

The mark does not have to be in quarters. Half note = 138, the marking on the first movement of Beethoven's Hammerklavier sonata, means 138 half notes per minute, so a half note is 60/138=0.435 s and a quarter note is 0.217 s. Converting to quarters first, then applying the formula, keeps this from going wrong.

The device that made such marks possible is recent. Johann Maelzel patented a workable metronome in 1815, and Beethoven published metronome marks for his symphonies in 1817, the first major composer to do so. Before that a tempo instruction was a word like allegro, which is a character more than a speed, and the difference between the two kinds of instruction has never been fully settled: the Hammerklavier marking is still argued about, since at 138 the movement is close to unplayable.

Example. At a tempo of quarter note = 92, how long does a dotted half note last?

A dotted half is 3 quarter notes, and each quarter is 60/92=0.6522 s, so t=3×0.6522=1.957 s.

Now you. At quarter note = 132, how long does a passage of five dotted quarter notes last?

Answer

A dotted quarter is 1.5 quarters, so it lasts 1.5×60/132=0.6818 s, and five of them last 5×0.6818=3.409 s.

The same arithmetic gives a whole movement a running time. Thirty-two bars of four quarter notes each, at quarter note = 120, contain 32×4=128 quarter notes at 0.5 s apiece, so 64 s. That is why a metronome mark and a bar count are enough to know whether a piece will fit on one side of a record, and why publishers have cared about them since long before anyone could measure a performance.

What the written durations do not say

Notated rhythm is a good approximation of performed rhythm and not a transcription of it, and three gaps are worth naming.

The first is that performers stretch time on purpose. Rubato means exactly that, taking from one note and giving to another, and in most repertoire a phrase slows slightly at its end whether or not anything is written. A performance that matched the arithmetic exactly would sound mechanical, which is the usual complaint about a first pass at engraving software playback.

The second is that some traditions systematically play what is written unequally. Swung eighth notes in jazz are written as plain pairs and played closer to long-short in a ratio near 2:1, so at quarter note = 120 the pair occupies 0.5 s split roughly 0.33 s and 0.17 s rather than 0.25 s each. French baroque notes inégales did the same thing two centuries earlier. The convention lives outside the notation entirely, and a reader who does not know the style will read the page correctly and play it wrongly.

The third is that the tree assumes division by two, patched by dots for three. Genuine division of a beat into three, five or seven equal parts needs a further device, the tuplet, which belongs with the next lesson because it depends on what a beat is.

And that is the gap this lesson leaves open. Every duration on the page is now measurable, and nothing yet says which notes are strong. A row of eight identical eighth notes could be four groups of two or two groups of four, and those are different pieces of music. Supplying the grouping is what metre does.

Metre and time signatures

Eight eighth notes in a row can be four groups of two or two groups of four, and those are different pieces of music. The previous lesson gave every note a length; this one gives the lengths a shape.

Pulse, beat and bar

Listen to almost any music and you will find yourself tapping. What you are tapping is the beat, a regular pulse the music implies, and what makes it feel regular in a particular way is that some beats feel stronger than others. Group the beats into a repeating pattern of strong and weak and you have metre.

Notation writes the grouping with barlines, vertical strokes cutting the staff into bars (American: measures). The rule is simple and almost never broken: the beat immediately after a barline is the strongest in the bar. It is called the downbeat, from the conductor's arm, and everything else in the bar is heard relative to it.

Be clear about what "strong" means, because it is not the same as "loud". A strong beat is a position in a pattern, and the pattern is maintained by the listener even when the performer does nothing to mark it. Play a bar of silence in the middle of a waltz and the listener still knows where the next downbeat falls. This is why metre can be contradicted, at the end of this lesson, and still be felt.

What the two numbers mean

The time signature at the head of the piece declares the pattern. It is two numbers stacked without a line between them, and the safest reading is: the bottom number names a note value, the top number says how many of them fill a bar.

The bottom number names the value as a fraction of a whole note. A 4 means quarter notes, an 8 means eighths, a 2 means halves, a 16 means sixteenths. So 3/4 is three quarter notes per bar, 2/2 is two half notes per bar, and 6/8 is six eighth notes per bar. A bar therefore lasts top divided by bottom whole notes: 3/4 gives three quarters of a whole note, 6/8 gives six eighths, which is also three quarters. The two signatures hold identical amounts of time, and, as the next section shows, they are still not the same metre.

Two signatures are printed as letters for historical reasons. C means 4/4, and it is not an abbreviation of "common"; it is the remains of a medieval circle-and-semicircle system in which a full circle meant triple division and a broken one duple. C with a vertical stroke through it means 2/2, called cut common or alla breve.

Example. How long does one bar of 3/4 last at a tempo of quarter note = 144?

Each quarter lasts 60/144=0.4167 s and the bar holds three of them, so 3×0.4167=1.250 s.

Now you. How long does one bar of 2/2 last at a tempo of half note = 60, and how many quarter notes does the bar contain?

Answer

A half note lasts 60/60=1 s and the bar holds two, so the bar lasts 2 s. It contains four quarter notes, exactly as a bar of 4/4 would, at the same total length.

Simple and compound, or how the beat divides

Here is the distinction that trips up nearly everyone, and it is worth getting right once. The top number does not always give the number of beats. It gives the number of note values, and in some metres those values are divisions rather than beats.

In simple metres the beat is an undotted note and divides into two. 2/4, 3/4 and 4/4 are simple, with two, three and four quarter-note beats, each splitting into two eighths.

In compound metres the beat is a dotted note and divides into three. 6/8 is the important case: it holds six eighth notes, but they are felt as two beats of three, so the beat is a dotted quarter and there are two of them. Counting six in a bar of 6/8 is the standard beginner's error, and it turns a lively 6/8 jig into a slow six. Similarly 9/8 is three dotted-quarter beats and 12/8 is four.

The test is arithmetic. If the top number is 6, 9 or 12, divide it by 3 to get the number of beats, and the beat is a dotted version of the value named by the bottom number. Otherwise the top number is the number of beats already.

That gives the answer to 3/4 against 6/8. Both bars contain six eighth notes. In 3/4 they group as 2+2+2, three beats each divided in two. In 6/8 they group as 3+3, two beats each divided in three. Leonard Bernstein's America famously alternates the two bar by bar, and the whole effect of the song lives in that difference.

Example. A piece is in 9/8. How many beats are in a bar, what note gets the beat, and how long is a bar at dotted quarter note = 80?

Nine divided by three is three beats, each a dotted quarter. At 80 dotted quarters per minute each beat lasts 60/80=0.75 s, so the bar lasts 3×0.75=2.25 s.

Now you. A piece is in 12/8 at dotted quarter note = 80. How many beats per bar, and how long is one eighth note?

Answer

Twelve divided by three is four dotted-quarter beats, so the bar lasts 4×0.75=3.00 s. Each beat divides into three eighths, so an eighth note lasts 0.75/3=0.25 s.

The accent hierarchy

Within a bar the strengths are graded rather than binary. In 4/4 the pattern is strong, weak, medium, weak: beat 1 is the downbeat, beat 3 carries a secondary accent, and beats 2 and 4 are weak. That is precisely why 4/4 is not two bars of 2/4 written lazily. In 2/4 every second beat is a downbeat; in 4/4 every fourth is, and beat 3 is audibly subordinate to beat 1. In 3/4 the pattern is strong, weak, weak, with no secondary accent, which is what makes triple metre feel unlike anything duple.

The hierarchy continues below the beat. Count a bar of 4/4 in eighths as "one and two and three and four and": the numbers are stronger than the "and"s, and the beat numbers are themselves graded as above. A note landing on an "and" is metrically weaker than one landing on a number, however loudly it is played, and that is the raw material for both syncopation and, much later in this course, the rules about where a dissonance may fall.

Beaming shows the reader the pattern

Eighth notes and shorter can have their flags replaced by beams, thick lines joining the stems. Beams are not decoration and are not free: they are drawn so that each beamed group fills one beat.

In 4/4, eighth notes beam in twos, one group per quarter-note beat, or in fours across a half bar. In 6/8 they beam in threes, one group per dotted-quarter beat. A page of 6/8 therefore looks different from a page of 3/4 even before you read the time signature, and that is the point: beaming lets a player see the metre rather than deduce it. Beaming across the middle of a bar of 4/4, joining the second and third beats into one group, is considered an engraving error precisely because it hides where beat 3 is.

The one traditional exception is vocal music, where beams were historically drawn to match syllables, one beam group per syllable, so a singer could see the word-setting. Modern editions increasingly beam vocal music by beat like everything else, and reprints of older editions do not.

Anacrusis, tuplets, and the two ways to leave the pattern

A piece need not start on a downbeat. An anacrusis, or pickup, is one or more notes before the first full bar, written in a short bar of their own. The convention, in older engraving, is that the final bar of the piece is shortened by the same amount so the two incomplete bars add to one full one, and Happy Birthday, in 3/4, is the example everybody has already sung: the two notes of "hap-py" are the pickup, and the downbeat of the first full bar lands on "birth".

The other departure is dividing a beat by something the notation cannot express. A tuplet does it by fiat: write the notes, bracket them, and print the number of them that are to fit in the time normally taken by a different number. A triplet is three notes in the time of two, marked with a 3, and it is how a simple metre borrows the compound division. In 4/4 at quarter note = 120, a quarter-note beat lasts 0.5 s, so each note of an eighth-note triplet lasts 0.5/3=0.1667 s, against 0.25 s for an ordinary eighth. Other ratios are written the same way, so a 5 over a group of sixteenths means five in the time of four. Compound metres borrow in the other direction with a duplet, two notes in the time of three.

Example. In 6/8 at dotted quarter note = 60, how long is one eighth note, and how long is each note of a duplet written across one beat?

The beat lasts 60/60=1 s and divides into three eighths, so an eighth is 1/3=0.333 s. A duplet fits two notes into that same 1 s beat, so each lasts 0.5 s.

Now you. In 4/4 at quarter note = 100, a beat is filled by a five-note tuplet of sixteenths, marked 5 in the time of 4. How long is each of those notes?

Answer

The beat lasts 60/100=0.6 s and is shared equally by five notes, so each lasts 0.6/5=0.12 s, against 0.6/4=0.15 s for an ordinary sixteenth.

Syncopation

Syncopation is an accent where the metre says there should not be one, and it only exists because the metre is strong enough to be contradicted. There are three standard ways to produce it and they are often combined.

The first is the tie. Sound a note on the weak half of beat 2 and tie it over beat 3, and beat 3 arrives with nothing articulated on it, so the ear promotes the note that came early. The second is the rest: put silence on the downbeat and the next attack, wherever it falls, takes the weight. The third is the printed accent mark, which simply insists.

The classic pattern in popular music is the Charleston figure: in a bar of 4/4, a dotted quarter, another dotted quarter, then a quarter. The attacks fall on beat 1, on the second half of beat 2, and on beat 4, so only the first coincides with a beat and the middle one falls in the metrically weakest position available. Backbeat, the rock and pop habit of putting the drum on beats 2 and 4, is the same idea made permanent: the weak beats are the loud ones, and the metre survives because everything else in the texture still points at 1 and 3.

Where metre stops working

Three honest limits. First, metre is a construction in the listener, not a property of the sound, and ambiguous passages exist where two hearings are both defensible until a cadence settles it. Second, plenty of music is not metrical at all: Gregorian chant, much unaccompanied folk singing, and the introductory alap of a raga are all measured in phrases rather than bars. Third, the simple-versus-compound scheme is a Western generalisation. Balkan dance music routinely uses additive metres in which the bar is built from unequal groups, a 7/8 grouped 2+2+3 rather than into seven equal beats, and the time signature alone will not tell you which grouping is meant, so good editions print it, as 3+2+2 over 8.

Rhythm is now fully readable: every note has a length, and every length has a place in a pattern. Pitch is still only a list of names with nothing said about the distances between them, and that is the subject of the next lesson.

Intervals

Two notes at once, or one after the other, make a distance, and naming that distance is the single most useful skill in reading a score. This lesson builds the naming system, which needs two numbers rather than one, and shows where the names come from.

Why one number is not enough

Play C4 then E4, then play C4 then E flat. Both pairs cover three letter names, C to E, and both are called some kind of third. They do not sound the same: the first is four semitones wide and bright, the second is three semitones wide and dark, and the whole character of a piece can turn on which one is used.

So an interval name has two parts. The number counts letter names, ignoring sharps and flats entirely. The quality counts semitones, and separates the intervals that share a number. C to E is a major third; C to E flat is a minor third. Neither part alone identifies the interval, and neither is redundant: the number tells you what the notes look like on the staff, and the quality tells you what they sound like.

Counting the number

The count is inclusive, meaning both end notes are counted. C up to E covers C, D, E, so it is a third. C up to G covers C, D, E, F, G, so a fifth. C up to the C above covers eight letters and is an eighth, which is the octave, and octave is simply the Latin for eighth. A note with itself is a first, called a unison.

Inclusive counting has one consequence that surprises everyone once and then never again: interval numbers do not add. Stack a third on a third and you do not get a sixth. C up to E is a third, E up to G is a third, and C up to G is a fifth, because the note E was counted twice, once as the top of the lower interval and once as the bottom of the upper. The rule is a+b-1: two thirds make 3+3-1=5, a third plus a fourth makes a sixth, and two fifths make a ninth.

Accidentals never change the number. F to A is a third and so are F sharp to A, F to A flat, and F sharp to A sharp, because all four span three letters. Only the quality moves.

Quality, and where the two families come from

Rather than memorising a table, generate it. Take the lower note of the interval, build the major scale on it (the next lesson derives that scale properly; for now, C major is the white keys), and look for the upper note. If the upper note belongs to that scale, the interval is perfect when the number is 1, 4, 5 or 8, and major for 2, 3, 6 and 7.

Applying that to C, the intervals up to each white key are: unison 0 semitones, major second 2, major third 4, perfect fourth 5, perfect fifth 7, major sixth 9, major seventh 11, octave 12. That is the full reference table, and it is worth knowing cold.

The split into two families is not arbitrary bookkeeping. The perfect intervals are the ones that appear earliest in the harmonic series and sit at the simplest ratios: 1:1, 4:3, 3:2 and 2:1. They also have the property that inverting one gives another perfect interval, which the major and minor pairs do not. Medieval theory treated exactly these as stable, which is why they inherited the word perfect, and their names have outlived the theory.

From the two families, four more qualities are reachable by moving one semitone at a time. Lower a major interval by a semitone and it becomes minor. Lower a minor interval further, or a perfect interval, by a semitone and it becomes diminished. Raise a major or a perfect interval by a semitone and it becomes augmented. The chains run

diminishedminormajoraugmented

for seconds, thirds, sixths and sevenths, and

diminishedperfectaugmented

for unisons, fourths, fifths and octaves. There is no such thing as a minor fifth or a major fourth, and using those names is the fastest way to tell a reader you have not done this lesson.

Example. Name the interval from E flat 4 up to C5.

Count letters first: E, F, G, A, B, C is six, so it is a sixth. Now count semitones: E flat is 3 semitones above C and the C above is 12, so the distance is 12-3=9 semitones. A major sixth is 9 semitones. So it is a major sixth.

Now you. Name the interval from A3 up to F4.

Answer

A, B, C, D, E, F is six letters, so a sixth. A is 9 semitones above C, and F above it is 17 semitones above that same C, so the interval is 8 semitones. A major sixth is 9, so 8 is one less: a minor sixth.

Building an interval upward

The reverse operation, given a note and an interval name, is the one that appears in every harmony exercise, and doing it in the right order prevents almost all mistakes. Get the letter first, from the number, and then adjust with accidentals until the semitone count is right. Never pick the pitch first and then look for a spelling, or you will write a diminished fourth where an augmented third was wanted.

Example. Write the note a minor sixth above F sharp 4.

The number is 6, so counting six letters from F gives F, G, A, B, C, D: the letter is D, and no other letter is permitted. A minor sixth is 8 semitones. F sharp 4 is 6 semitones above C4, so the target is 14 above C4, which is 2 above C5, which is D5 natural. The answer is D5, and it checks: F sharp to D is a sixth of 8 semitones.

Now you. Write the note an augmented fourth above B flat 3.

Answer

Four letters from B gives B, C, D, E, so the letter is E. An augmented fourth is 6 semitones, since a perfect fourth is 5. B flat 3 is 10 semitones above C3, so the target is 16 above C3, which is 4 above C4: E4 natural. The answer is E4, not F4, even though F4 sounds nearer to what an untrained eye expects.

The tritone

One interval deserves its own paragraph. Six semitones is exactly half of twelve, and it is the only interval that inverts into itself. It arises as an augmented fourth, from F up to B, and as a diminished fifth, from B up to F, which are the same six semitones spelled two ways, and the neutral name for both is the tritone, so called because it spans three whole tones.

In equal temperament its ratio is 26/12=2=1.4142, which is irrational, so it is as far from a simple whole-number ratio as an interval can get. The nearest small ratios are 45:32 at 590.2 cents and 7:5 at 582.5 cents, neither of them close and neither of them simple. Counterpoint treatises restricted it from the Middle Ages onward, and it later picked up the nickname diabolus in musica, the devil in music; nineteenth-century composers then used it for exactly the reason it had been restricted. When progressions arrive it turns out to be the engine inside the dominant seventh chord, which is where most of the harmonic motion in tonal music comes from.

Inversion halves the work

Move the lower note of an interval up an octave, or the upper note down one, and you have its inversion. C up to E, a major third, becomes E up to C, a sixth. Two rules cover every case.

The numbers add to 9. This falls straight out of inclusive counting: the two intervals together span an octave, whose eight letters are counted twice at the shared note, so 8+1=9. A third inverts to a sixth, a second to a seventh, a fourth to a fifth, a unison to an octave.

The quality flips: major becomes minor, minor becomes major, augmented becomes diminished, diminished becomes augmented, and perfect stays perfect. Check it in semitones, which must add to 12: a major third is 4, a minor sixth is 8, and 4+8=12. A perfect fifth is 7 and a perfect fourth is 5, so the perfect family maps onto itself, which is the property the word perfect was pointing at.

Inversion is worth using because it halves what has to be recognised at speed. Sevenths and sixths are hard to judge by eye on the staff; thirds and seconds are easy. Invert mentally, name the easy one, convert.

Example. What is the inversion of a diminished fifth, and does the semitone arithmetic agree?

9-5=4, so it is some kind of fourth, and diminished becomes augmented: an augmented fourth. A diminished fifth is 6 semitones and an augmented fourth is 6, and 6+6=12. The tritone inverting into itself, as promised.

Now you. What is the inversion of a minor seventh, and what are the two semitone counts?

Answer

9-7=2, and minor becomes major: a major second. A minor seventh is 10 semitones, a major second is 2, and 10+2=12.

Bigger than an octave

Intervals wider than an octave are compound. Their names continue the count, so an octave plus a second is a ninth, an octave plus a third is a tenth, an octave plus a fifth is a twelfth. To get the simple interval inside, subtract 7 rather than 8, because of inclusive counting again: a ninth reduces to a second, an eleventh to a fourth, a thirteenth to a sixth.

Quality carries over unchanged, so a major tenth is a major third plus an octave. In practice only ninths, elevenths and thirteenths are named as compounds, because they appear as chord extensions; anything else is usually described as "a tenth" or just as its simple form with the octave taken for granted.

Consonance, and what the ear is actually doing

The first lesson showed that simple frequency ratios share upper partials. Now the intervals have names, the list can be written down properly. The perfect consonances are the octave at 2:1, the fifth at 3:2 and the fourth at 4:3. The imperfect consonances are the major third at 5:4, the minor third at 6:5, the major sixth at 5:3 and the minor sixth at 8:5. Everything else, the seconds, the sevenths and the tritone, is treated as dissonant, and the ratios show why: a major second is 9:8 and a major seventh 15:8, which share almost nothing in the audible part of the spectrum.

Two honest qualifications. First, there is a physical mechanism underneath, and it is not the arithmetic itself: partials that fall close together in frequency, within roughly a critical band of the ear, beat against each other and produce the sensation called roughness. Simple ratios avoid this because partials either coincide exactly or separate cleanly, and that is why the perceptual ranking tracks the numerical one.

Second, the categories are historical as much as acoustic, and the fourth is the proof. It has the third-simplest ratio of all, yet counterpoint from the fifteenth century onward treats a fourth above the bass as a dissonance requiring resolution, while a fourth between upper voices is fine. No acoustic account will produce that rule, because it is not an acoustic fact; it is a stylistic one, in a style that thought in terms of the bass. Expect the ratios to explain the broad ranking and expect the fine print to be cultural.

Intervals name distances between any two notes at all. Music does not use any two notes: it draws from a fixed selection, seven of the twelve, chosen so that the intervals inside it come out a particular way. That selection is the scale, and it is next.

Scales and degrees

Music almost never uses all twelve pitches equally. It picks a small set, usually seven, and treats one of them as home, and this lesson is about which seven and why the choice has the consequences it does.

The major scale is a pattern of steps

Play the white keys from C up to the next C and you have the major scale, the reference object for everything that follows. What matters is not the letters but the pattern of distances. Measured in semitones, C to D is 2, D to E is 2, E to F is 1, F to G is 2, G to A is 2, A to B is 2, B to C is 1. Writing a whole tone as T and a semitone as S, the pattern is

T T S T T S T

and the seven steps sum to 2+2+1+2+2+2+1=12, which is the octave, as any scale must.

The pattern is the definition, not the white keys. C major happens to use white keys because the keyboard was laid out around it, and any other starting note gives the same scale transposed, provided the pattern is preserved exactly. Start on G and apply T T S T T S T: G, A, B, C, D, E, and then the sixth step of a tone from E lands on F sharp, and a final semitone reaches G. So G major is G A B C D E F# G, and the sharp is not a decoration or a stylistic choice. Use F natural and the last two steps become a semitone and a tone rather than a tone and a semitone, which is a different scale.

One letter each, no exceptions

Notice that G major used every letter name once, in order, and that this is what forced the spelling F sharp rather than G flat. The rule holds for every major and minor scale: seven notes, seven consecutive letters, each appearing exactly once, with accidentals doing whatever they must.

The rule earns its keep on the flat side. Build a major scale on F: F, G, A, then a semitone to B flat, then C, D, E, F. The fourth note is a black key, and calling it A sharp would be wrong even though it is the same key on a piano, because the scale would then read F G A A# C D E and use A twice while skipping B. Spelled B flat, the scale is legible on the staff as one notehead per position, which is the whole reason the convention exists.

The extreme cases follow the rule without flinching. C sharp major is C# D# E# F# G# A# B# C#, with an E sharp and a B sharp that sound like F and C, because the letters E and B have to appear and be raised. C flat major is Cb Db Eb Fb Gb Ab Bb Cb, with an F flat. These look absurd and are perfectly correct, and the next lesson explains when a musician would rather use them than the enharmonic alternative.

Example. Write the D major scale.

Apply T T S T T S T from D. D up a tone is E, up a tone is F sharp, up a semitone is G, up a tone is A, up a tone is B, up a tone is C sharp, up a semitone is D. So D E F# G A B C# D, with each letter used once.

Now you. Write the E flat major scale.

Answer

Eb F G Ab Bb C D Eb. Checking the steps: Eb to F is 2, F to G is 2, G to Ab is 1, Ab to Bb is 2, Bb to C is 2, C to D is 2, D to Eb is 1, which is T T S T T S T.

The degrees are not interchangeable

The seven notes of a scale are its degrees, numbered 1 to 7 from the tonic, and the numbering is the most portable idea in the subject: a statement about degree 5 is true in every key at once, which is why professional discussion of harmony is conducted almost entirely in numbers.

Each degree also has a name, and the names encode a structure worth seeing rather than memorising. Degree 1 is the tonic, home. Degree 5 is the dominant, a perfect fifth above the tonic, the strongest relation there is after the octave. Degree 4 is the subdominant, and the "sub" is not "one below the dominant": it is the note a perfect fifth below the tonic, which lands on degree 4 an octave up. So the tonic sits between two fifths, one on each side, and the two nearest relatives are named for that.

The remaining names extend the scheme. Degree 3 is the mediant because it lies midway between tonic and dominant; degree 6 is the submediant, midway between tonic and subdominant on the way down. Degree 2 is the supertonic, simply the note above the tonic. Degree 7 is the leading note, or leading tone, and its name is a claim about behaviour that the next section justifies.

Example. In E flat major, name degrees 4, 5 and 7.

The scale is Eb F G Ab Bb C D. Degree 4 is A flat, the subdominant; degree 5 is B flat, the dominant; degree 7 is D, the leading note.

Now you. In B major, name the dominant, the subdominant and the leading note.

Answer

B major is B C# D# E F# G# A#. The dominant is F sharp, the subdominant is E, and the leading note is A sharp.

Why degree 7 pulls

Look at where the two semitones fall in T T S T T S T. One is between degrees 3 and 4, the other between degrees 7 and 8, that is, between the leading note and the tonic an octave up. Everything characteristic of major tonality comes from those two positions.

The semitone from 7 to 8 is the engine. A leading note sounds unfinished: play C D E F G A B and stop, and almost every listener raised on this music wants the C. The best current explanation is expectation learned from exposure rather than anything acoustic, since a semitone is the smallest available step and the note above it is the most frequent note in the style, but whatever the mechanism, the effect is reliable and composers have relied on it for four hundred years. When chords arrive, that single semitone will turn out to be half of what makes a dominant chord resolve.

The consequence to hold on to is that a scale is not a set of pitches. It is a set of pitches plus a hierarchy, and moving the same seven notes around a different centre gives different music, which is the point of the final section of this lesson.

Minor comes in three versions

The natural minor scale has the step pattern

T S T T S T T

which on A gives the white keys again: A B C D E F G A. Compared with A major it flattens degrees 3, 6 and 7, and the flattened third is most of what makes it sound the way it does.

There is a problem, and it is precisely the leading note. In A natural minor, degree 7 is G, a whole tone below the tonic rather than a semitone, so it does not pull. Composers writing minor-key music in the tonal style wanted that pull, and bought it back by raising degree 7, giving the harmonic minor: A B C D E F G# A, with steps T S T T S A2 S where that A2 is three semitones.

The cost shows up immediately. Raising G to G sharp opens a gap of three semitones between F and G sharp, an interval of an augmented second, since it spans two letters. It is awkward to sing and instantly recognisable, and it is why harmonic minor has a colouring that most listeners associate with music from far outside the classical tradition even though it was invented inside it.

The melodic minor is the patch for the patch. Raise degree 6 as well and the gap closes: A B C D E F# G# A, which is a smooth line with a working leading note. Classical practice raises 6 and 7 only when ascending and uses the natural minor coming down, on the argument that the leading note is only needed when the line is heading for the tonic. Jazz practice keeps the raised form in both directions and calls it the melodic minor scale full stop. Both are in current use, which is a good reminder that these three scales are not three facts about nature but three solutions to one compositional problem.

Example. Write C harmonic minor and identify the augmented second.

C natural minor is C D Eb F G Ab Bb C. Raise degree 7 from B flat to B natural: C D Eb F G Ab B C. The augmented second is A flat up to B natural, three semitones across two letters.

Now you. Write F sharp harmonic minor.

Answer

F sharp natural minor is F# G# A B C# D E F#. Raising degree 7 gives F# G# A B C# D E# F#. The seventh must be spelled E sharp, not F natural, since the letter E has to appear once. The augmented second is D to E sharp.

Scales with fewer notes, and with more

Not every scale has seven degrees. The major pentatonic takes degrees 1, 2, 3, 5 and 6, dropping exactly the two notes that were adjacent to a semitone, so C major pentatonic is C D E G A. With no semitones left, no two notes in it clash, which is why it is nearly impossible to play something ugly on it. This is also why the five black keys work as a group: F# G# A# C# D# is the pentatonic of G flat major, and improvising on the black keys alone is the standard party demonstration of that fact.

The minor pentatonic is its rotation, taking degrees 1, flat 3, 4, 5 and flat 7, so A minor pentatonic is A C D E G, the same five pitches as C major pentatonic around a different home. Add a flattened fifth as a passing note, E flat here, and you have the blues scale.

In the other direction, the chromatic scale uses all twelve semitones and so has no hierarchy at all; it is a supply of notes rather than a key, and appears inside tonal music as decoration. The whole tone scale takes six equal whole-tone steps, C D E F# G# A#, and being perfectly symmetrical it has no semitone, no leading note and no obvious tonic. Debussy used it for exactly that quality of floating, and there are only two distinct whole tone scales, since the third one you try turns out to be the first transposed.

The modes are rotations, but a rotation is not enough

Play the white keys from D to D and the step pattern is T S T T T S T. That is not major and not natural minor; it is a third arrangement, called Dorian. Rotating the major scale to start on each of its degrees gives seven such patterns: Ionian on 1 (major), Dorian on 2, Phrygian on 3, Lydian on 4, Mixolydian on 5, Aeolian on 6 (natural minor), Locrian on 7.

The useful way to hold them is as alterations of a scale you already know rather than as rotations you have to count. Against major, Lydian raises degree 4 and Mixolydian flattens degree 7. Against natural minor, Dorian raises degree 6 and Phrygian flattens degree 2. Locrian flattens degree 5 as well as 2, 3, 6 and 7, which leaves its tonic triad diminished and is why it is nearly unusable as a key.

The honest limit is the one people skate over. Playing the white keys from D does not make music Dorian, because a set of pitches carries no centre with it. Whether a passage is in D Dorian or in C major depends entirely on what the music does: which note the phrases land on, which chord ends them, which note the bass insists on. Modal writing gets its centre by repetition and by drone or pedal rather than by the leading-note pull that major and minor use, and a modal passage with a strong V chord in it usually stops sounding modal and starts sounding like ordinary minor with a decorated sixth.

Writing any of these out in full is tedious, because the accidentals have to be printed on every affected note in every bar. The shorthand that fixes this is the key signature, and the order in which its sharps and flats appear turns out not to be arbitrary at all.

Keys and key signatures

A piece in E major contains four notes that are always sharp, and printing the sign on every one of them would bury the page in accidentals. The key signature is the fix, and its layout is not a convention anyone chose freely.

What a key signature declares

A key signature is a group of sharps or flats printed immediately after the clef, at the start of every staff line. It says that those notes are sharp or flat throughout the piece, until it is cancelled or replaced.

Its scope is wider than an accidental's in one important way. An accidental applies to one letter in one octave for the rest of one bar; a key signature applies to that letter in every octave, in every bar, until the end. A signature with an F sharp in it sharpens F2, F4 and F6 alike, even though the sign is printed on one line only. An accidental inside the piece temporarily overrides the signature for its own bar and octave, and the barline hands control back.

A key is a scale plus the tonic that scale is built on, and the signature exists to declare which. That declaration is genuinely ambiguous, as the last section of this lesson shows, but for now: no sharps or flats means C major, one sharp means G major, two means D major, and there is a rule behind that sequence.

The order of the sharps is forced

Build the major scales in order and watch what happens. C major needs nothing. Go up a perfect fifth to G major, and the scale G A B C D E F# G needs exactly one accidental, F sharp, because degree 7 of a major scale must be a semitone below the tonic. Go up another fifth to D major: D E F# G A B C# D, which keeps the F sharp and adds C sharp, again as its leading note. Another fifth to A major adds G sharp. Another to E major adds D sharp.

Two things are happening at once and both are worth stating. First, moving up a fifth adds exactly one sharp and never removes one, because the new scale shares six of its seven notes with the old and differs only in raising the old degree 4. Second, the new sharp is always the leading note of the new key, which is a fifth above the previous new sharp, since the keys themselves are a fifth apart.

So the order in which sharps appear is itself a chain of fifths:

FCGDAEB

F sharp, C sharp, G sharp, D sharp, A sharp, E sharp, B sharp. They are always printed in that order and a signature never skips one, so four sharps is always F# C# G# D# and never any other four.

Going the other way, down a fifth from C to F, gives F G A Bb C D E F, which needs B flat, since degree 4 of the new key must be a semitone above degree 3. Down another fifth to B flat major adds E flat. The order of flats is a chain of falling fifths, which is the sharp order read backwards:

BEADGCF

That the two lists are exact reverses of one another is not a coincidence to be memorised. It follows from the fact that going up a fifth and going down a fifth are inverse operations, and the printed signatures inherit the symmetry.

Reading a key off a signature

Two rules, each a direct consequence of the derivation above.

For sharps: the last sharp in the signature is the leading note, so the tonic is one semitone above it. Four sharps ends on D sharp, and a semitone above D sharp is E, so it is E major.

For flats: the last flat is degree 4 of the key, and counting back four degrees is fiddly, so use the shortcut instead. In any signature of two or more flats, the second-to-last flat is the tonic. Five flats is Bb Eb Ab Db Gb, the second-to-last is D flat, and the key is D flat major. One flat is the single exception, since there is no second-to-last, and it is F major.

Here is the whole system in one place, with each major key beside the minor key that shares its signature.

SharpsMajorMinorFlatsMajorMinor
0CA0CA
1GE1FD
2DB2BbG
3AF#3EbC
4EC#4AbF
5BG#5DbBb
6F#D#6GbEb
7C#A#7CbAb

Example. A piece has a signature of four sharps and ends on a C sharp minor chord. What key is it in, and what are the four sharps?

The four sharps are F#, C#, G# and D#, in that order. That signature is E major or C sharp minor, and the final chord settles it: C sharp minor.

Now you. A piece has five flats. Name the flats in order, and both keys the signature could mean.

Answer

Bb, Eb, Ab, Db, Gb. The second-to-last flat is D flat, so the major key is D flat, and its relative minor is B flat minor.

The circle of fifths, and why it closes

Chain the keys by fifths and something convenient happens. C, G, D, A, E, B, and then F sharp, C sharp: eight keys, with sharps accumulating one at a time. Chain downward instead and you get C, F, B flat, E flat, A flat, D flat, G flat, C flat, with flats accumulating. Twelve steps of a fifth from C arrive back at C, which lets the whole system be drawn as a clock face with C at twelve, the sharp keys clockwise and the flat keys anticlockwise. That drawing is the circle of fifths, and it is the single most useful diagram in the subject: neighbours on it share six of seven notes, and distance around it is a good proxy for how far apart two keys sound.

It closes only because of a compromise made in the first lesson. Twelve pure fifths give a ratio of (3/2)12=129.746 while seven octaves give 27=128, so a chain of pure fifths overshoots by the Pythagorean comma of 23.5 cents and never returns to its starting pitch. Equal temperament shaves each fifth by 1.96 cents, twelve of which is exactly the comma, and the circle snaps shut. In the unequal temperaments of the seventeenth century the far side of the circle was genuinely unusable, and the modern freedom to modulate anywhere was bought with those two cents per fifth.

Enharmonic keys, and choosing between them

Because the circle closes, the sharp end and the flat end meet, and three keys have two spellings. B major with five sharps is the same set of pitches as C flat major with seven flats. F sharp major with six sharps equals G flat major with six flats. C sharp major with seven sharps equals D flat major with five flats.

The choice is normally made on accidental count, which is why D flat major is common in print and C sharp major is not. It is not automatic, though. A movement that has just been in B major may write C sharp minor rather than D flat minor because the players are already reading sharps, and a horn or clarinet part transposed from a flat key stays in flats to keep its own fingerings sane. The pitches are identical in equal temperament; the page is not, and the page is what a player reads at speed.

Example. Write the key signature of A major, and say which enharmonic key, if any, could replace it.

A major has three sharps: F#, C# and G#. Its enharmonic equivalent would be B double flat major, which is not a key anybody writes, so there is no alternative. Only the keys near the six-accidental crossing point have practical enharmonic twins.

Now you. Write the key signature of A flat major, and name its enharmonic equivalent if it has a usable one.

Answer

A flat major has four flats: Bb, Eb, Ab and Db. Its enharmonic equivalent is G sharp major, whose scale is G# A# B# C# D# E# F double sharp, so its signature would need a double sharp and is never printed. A flat is far from the crossing point, exactly as A major was.

Relative and parallel

Two minor keys sit near every major one and they are easy to confuse.

The relative minor shares the signature. It is built on degree 6 of the major scale, which is a minor third below the tonic, so the relative minor of C major is A minor, of E flat major is C minor, of B major is G sharp minor. Both keys use the same seven pitches; they differ only in which is home, and that difference is audible enough to carry a whole movement.

The parallel minor shares the tonic instead. C major and C minor start on the same note, and their signatures differ by exactly three accidentals: C major has none, C minor has three flats. The three is not a coincidence, since natural minor flattens degrees 3, 6 and 7 of the major scale, and each of those flattenings is one step around the circle. E major has four sharps and E minor has one, again a difference of three.

Composers use the two relationships differently. A move to the relative minor is smooth, because nothing in the signature changes and the shared notes carry across. A switch to the parallel minor is a colour change, sudden and obvious, and is often used at the same pitch level for exactly that reason.

Example. Name the relative minor of E flat major and the parallel minor of E major, with the signature of each.

The relative minor of E flat major is C minor, sharing three flats. The parallel minor of E major is E minor, which has one sharp against E major's four.

Now you. Name the relative minor of B major and the parallel minor of B major, with the signature of each.

Answer

The relative minor is G sharp minor, sharing B major's five sharps. The parallel minor is B minor, with two sharps, three fewer than B major.

The signature does not settle the key

The last section is the caveat, and it is a real one. A signature is compatible with two keys by construction and with more than two in practice, so it is evidence rather than proof.

Three things settle it instead. The first is the ending: tonal pieces almost always finish on the tonic chord, so the last bass note is the strongest single clue, and the first bar is usually the second strongest. The second is the accidentals inside the piece. Minor-key writing raises degree 7 to get a leading note, and that raised note is not in the signature, so a piece with no sharps in its signature that keeps printing G sharp is in A minor rather than C major. Look for the raised seventh and the key names itself. The third is the harmony, which is what the rest of this course is for.

Two further complications are worth knowing about now. Music often changes key part way through, and if the new key is short-lived the signature does not change at all; the modulation shows up only as accidentals, and recognising that pattern is a skill the last lessons of this course build. And modal music sits badly in this system: a piece in D Dorian is often printed with no signature at all, or with one flat, and neither choice means what the table above says it means.

The signature has now given a fixed set of seven notes and a home among them. The next step is to stop treating them one at a time: stack them in thirds, and the result is chords.

Triads and sevenths

Everything so far has been one note at a time. A chord is three or more sounding together, and the whole of tonal harmony is built from one method of choosing which.

Why thirds

The method is to stack thirds: take a note, add the note two letters above it, then two letters above that. From C, the stack is C, E, G. Skip one letter each time and stop at three notes and you have a triad.

Thirds rather than any other interval, for a reason from the first lesson. The harmonic series of a low C runs 65.41, 130.81, 196.22, 261.63, 327.03, 392.44 Hz, and partials 4, 5 and 6 are 261.63, 327.03 and 392.44, in the ratio 4:5:6. That is a note, a major third above it, and a minor third above that, which is exactly the C major triad, sitting inside the spectrum of a single sounding C. The chord is not a human invention layered onto the physics; it is three of the loudest components of one note, made explicit.

Stacking seconds would give clusters of adjacent notes, which the previous lesson called dissonant. Stacking fourths gives a usable but unstable sound, used deliberately in twentieth-century music precisely because it is not this. Thirds give the two consonant intervals, 5:4 and 6:5, and their sum, the perfect fifth at 3:2, which is why a triad is stable.

The three notes have names: the root the stack starts on, the third, and the fifth, named by their interval above the root. The root is not necessarily the lowest sounding note, which the section on inversions makes precise.

Four qualities

Two thirds can each be major or minor, giving four combinations, and all four are used.

A major triad is a major third then a minor third, so 0, 4 and 7 semitones above the root. C E G. A minor triad is a minor third then a major third: 0, 3, 7. C Eb G. Both contain a perfect fifth, and both are consonant; the difference is only which third, and it is the difference between the two moods most listeners can name.

The other two stack identical thirds and lose the perfect fifth. A diminished triad is two minor thirds, 0, 3, 6, so its outer interval is the tritone: B D F. An augmented triad is two major thirds, 0, 4, 8, an augmented fifth: C E G#. Both are unstable and both are used for exactly that, the diminished constantly and the augmented sparingly.

Example. What quality is the triad F# A# C#?

F sharp to A sharp is 4 semitones, a major third. A sharp to C sharp is 3 semitones, a minor third. Major then minor is a major triad, so this is F sharp major.

Now you. What quality is the triad G Bb Db?

Answer

G to B flat is 3 semitones and B flat to D flat is 3 semitones, two minor thirds, so it is a diminished triad: G diminished. Its outer interval, G to D flat, is the tritone.

The seven chords of a key

Now restrict the stack to notes of one scale. Build a triad on each degree of C major, using only white keys, and every quality is decided for you by where the semitones fall.

On C: C E G, major. On D: D F A, minor, because D to F is only three semitones thanks to the E to F semitone. On E: E G B, minor. On F: F A C, major. On G: G B D, major. On A: A C E, minor. On B: B D F, diminished, since it contains both of the scale's semitone pairs.

The result is the same in every major key, because it depends only on the step pattern T T S T T S T, not on the letters. Major, minor, minor, major, major, minor, diminished, in that order, always.

These are written as roman numerals, one per degree, with case carrying the quality: uppercase for major, lowercase for minor, a small circle after a lowercase numeral for diminished and a plus sign after an uppercase one for augmented. So any major key has

IiiiiiIVVvivii

The numerals are the reason this course spends so long on scale degrees. They describe a progression without naming a key, so the same three symbols describe the same music in every key at once, and a musician who learns that ii goes well to V has learned it for all twelve.

Example. Spell the ii chord in A major and give its roman numeral and chord symbol.

A major is A B C# D E F# G#. Degree 2 is B, and the triad on it is B, D, F#. B to D is 3 semitones and D to F sharp is 4, so it is minor: the numeral is ii and the symbol is Bm.

Now you. Spell the chord on degree 7 of E flat major, and give its numeral.

Answer

E flat major is Eb F G Ab Bb C D. Degree 7 is D, and the triad is D, F, Ab. Both thirds are 3 semitones, so it is diminished: vii° in E flat, spelled D diminished.

Minor keys, and where the dominant comes from

Minor is less tidy, because the previous lesson gave it three forms. Using natural minor alone, A minor gives i, ii°, III, iv, v, VI, VII: the chord on degree 5 is E G B, which is minor.

That is a problem for the same reason the natural minor scale was a problem. A minor v has no leading note in it, so it does not pull back to the tonic. Raising degree 7 from G to G sharp, which is what harmonic minor exists to do, changes every chord that contains that degree: v becomes V, since E G# B is major; VII becomes vii°, since G# B D is diminished; and III becomes III+, since C E G# is augmented.

Practice takes what it wants from both forms. Real minor-key music uses i, ii°, iv, VI and the raised V and vii°, and the augmented III+ is a rarity. The working summary is that minor-key harmony is natural minor with a major dominant borrowed from the harmonic form, and the reason is entirely the leading note.

Inversions

A triad has three notes and any of them can be lowest. Root position puts the root in the bass. First inversion puts the third in the bass, so C E G becomes E G C. Second inversion puts the fifth in the bass: G C E.

The chord is the same chord in all three, since it holds the same three pitch classes and the same root, and it does not sound the same, because the bass line is what a listener tracks most strongly. First inversion is lighter and more mobile than root position and is used to keep a bass line stepping smoothly. Second inversion is genuinely unstable, since the interval from bass to root is a fourth, which the previous lesson noted is treated as a dissonance above a bass, and it is restricted to a few standard uses, the commonest being the cadential six-four in the next lesson.

Two notations exist. Figured bass, from the seventeenth century, writes intervals above the bass note: root position is 5/3 and normally left unmarked, first inversion is 6/3 abbreviated to 6, second inversion is 6/4. Combined with the numeral, that gives V6 for a first-inversion dominant and I6/4 for a second-inversion tonic. Chord symbols, from popular music, write a slash: C/E is a C chord with E in the bass. The two systems say the same thing, one relative to the key and one absolute.

One more third

Add a fourth note, another third above the fifth, and the triad becomes a seventh chord, named for the interval from root to top. There are five worth knowing, and they are distinguished by the triad underneath and the size of the seventh.

A major seventh is a major triad plus a major seventh, 0, 4, 7, 11 semitones, written Cmaj7. A dominant seventh is a major triad plus a minor seventh, 0, 4, 7, 10, written C7. A minor seventh is a minor triad plus a minor seventh, 0, 3, 7, 10, written Cm7. A half-diminished seventh is a diminished triad plus a minor seventh, 0, 3, 6, 10, written Cm7b5 or with a slashed circle. A fully diminished seventh is a diminished triad plus a diminished seventh, 0, 3, 6, 9, written Cdim7, and being three stacked minor thirds it is perfectly symmetrical, which is why it can resolve almost anywhere.

Doing the same restriction to one key gives the diatonic sevenths of C major: Cmaj7, Dm7, Em7, Fmaj7, G7, Am7 and Bm7b5, written Imaj7, ii7, iii7, IVmaj7, V7, vi7 and viiø7. Look at what the arithmetic has just delivered. Exactly one dominant seventh occurs naturally in a major key, and it is on degree 5. That is not a naming coincidence; it is why the chord is called what it is, and it is the single most important fact in the next lesson.

Sevenths invert too, with figures 7, 6/5, 4/3 and 4/2 for root position and the three inversions, so V4/2 is a dominant seventh with its seventh in the bass.

Example. Spell V7 in D major.

D major is D E F# G A B C#. Degree 5 is A, and the seventh chord on it stacks A, C#, E, G. Check the intervals: A to C sharp is 4, C sharp to E is 3, E to G is 3, so it is a major triad with a minor seventh, a dominant seventh. The symbol is A7.

Now you. Spell ii7 in B flat major and give its chord symbol.

Answer

B flat major is Bb C D Eb F G A. Degree 2 is C, and the stack is C, Eb, G, Bb: a minor triad with a minor seventh, so Cm7.

What a chord label leaves out

Three limits, all of which matter when analysing real music.

A chord symbol says which pitch classes, not how they are arranged. C E G, E G C an octave up, and a spread with three Cs in different octaves are all "C major", and they sound substantially different. Voicing, spacing and doubling are choices a label does not record, and the next lesson's voice-leading rules are largely about them.

A chord label also assumes there is a chord to label. Real music contains notes that belong to no chord at all, and the last lesson of this course deals with them; until then, be aware that picking the chord from a bar sometimes means deciding which notes are structural, and that two competent analysts occasionally disagree.

And the stacking-thirds scheme is a description of one repertoire, not of music in general. Jazz continues the stack to ninths, elevenths and thirteenths, and adds sixths and suspensions in which the third is replaced by a second or a fourth. Debussy and much film music use chords whose function this scheme cannot name. Rock and folk often use plain triads in progressions that break the rules the next lesson sets out, and they sound fine. The scheme is a very good account of European music from about 1600 to about 1900, and a partial account of everything since.

Chords now exist, but nothing so far says why one should follow another. A single triad, however well spelled, does nothing at all. What turns a list of chords into a progression is the subject of the next lesson.

How a progression works

A single chord does nothing. Harmony is motion, and this lesson is about which motions the tonal style uses and why those and not others.

Three jobs, not seven chords

The previous lesson produced seven triads per key. In practice they behave as three groups, and thinking in groups is what makes progressions predictable.

Tonic function means stability and arrival: I, and by sharing notes with it, vi and iii. Dominant function means maximum tension pointing at the tonic: V, and vii°, which is V7 with its root removed. Predominant function, also called subdominant, means a chord that leads to the dominant: IV and ii.

The standard flow is tonic, predominant, dominant, tonic. Moving forwards through that cycle sounds normal; skipping is fine, since I straight to V is everywhere; going backwards, V to IV, is the one move classical style avoids, and it is precisely the move blues and rock use constantly. That single disagreement is a useful early warning that these are conventions of one repertoire.

Underneath the grouping is a simpler fact about roots. The strongest root motion is falling by a perfect fifth, which is what V to I does, and ii to V, and vi to ii. Run the whole chain in C major and you get C, F, B diminished, E minor, A minor, D minor, G, C: the numerals I, IV, vii°, iii, vi, ii, V, I. Every adjacent pair falls a fifth, or rises a fourth, which is the same motion in a different octave, and the sequence is the skeleton of a great deal of Baroque music. Falling fifths feel strong for the reason given in the first lesson: the two roots share more partials than any other pair, so the ear hears a relation rather than a juxtaposition.

The engine inside V7

Take the dominant seventh of C major: G, B, D, F. Look at the outer two of those four notes, or rather at B and F. They are 6 semitones apart, a tritone, the interval the lesson on intervals described as the least stable in the system.

Now see what each note wants. B is degree 7, the leading note, a semitone below the tonic, so it pulls up to C. F is degree 4, a semitone above degree 3, and in this context it pulls down to E. Let both do what they want, and the two voices move outward by a semitone each in contrary motion, landing on C and E, which are the root and third of the tonic triad. Hold the G, move the D to C or E, and the chord is complete.

That is the whole mechanism. V7 resolving to I is one unstable interval collapsing to a stable one, with both moving voices travelling the smallest possible distance, and it is why the progression sounds not merely conventional but inevitable. Everything else in tonal harmony is arranged around delaying, decorating or evading that resolution.

Two consequences follow. First, the dominant seventh defines a key almost by itself: only one such chord occurs naturally in a major key, so hearing G7 tells you the key is C. Second, the definition is not quite unique, since the tritone B and F also lives inside D flat 7, whose root is a tritone from G. Jazz exploits exactly that, substituting Db7 for G7 and resolving to C anyway, and the substitution works because the crucial pair of notes is identical in both chords.

Example. Spell V7 in E flat major and resolve it, saying what each note does.

E flat major is Eb F G Ab Bb C D, so degree 5 is B flat and V7 is Bb, D, F, Ab. The tritone is D and A flat. D is the leading note and rises a semitone to E flat; A flat is degree 4 and falls a semitone to G; B flat is also the fifth of the tonic chord, so a voice holding it need not move at all; F falls to E flat. The tonic chord Eb G Bb arrives with nothing having moved more than a whole tone.

Now you. Spell V7 in A major and resolve it, naming the tritone and where each of its notes goes.

Answer

A major is A B C# D E F# G#, so V7 is E, G#, B, D. The tritone is G sharp against D. G sharp is the leading note and rises to A; D is degree 4 and falls to C sharp; the resolution is the tonic triad A C# E.

Cadences are the punctuation

A cadence is the harmonic formula that ends a phrase, and there are four to recognise.

The perfect cadence, or authentic cadence, is V to I. In its strongest form both chords are in root position and the melody ends on the tonic, which is a full stop. Weaken either condition, by inverting a chord or ending the melody on degree 3, and it becomes an imperfect authentic cadence, more like a semicolon.

The half cadence ends on V rather than moving to it. It is a question mark: the phrase stops on tension and the answering phrase resolves it. Any chord may precede it, and I to V or ii to V are the usual approaches.

The plagal cadence is IV to I, the "amen" at the end of a hymn. It contains no leading note, since IV holds degrees 4 and 6 but not 7, so it feels settled rather than driven and usually confirms an ending a perfect cadence has already made.

The interrupted cadence, also called deceptive, sets up V and then goes to vi instead of I. The bass rises a step where it was expected to fall a fifth, and the effect is of a phrase refusing to end, which is why it so often appears just before the real ending.

One decorated form is worth naming because it is everywhere and is easy to misread. The cadential six-four writes the tonic chord in second inversion on a strong beat over the dominant bass, then lets it resolve down to V before going to I: I6/4 to V to I. It looks like a tonic chord, but functionally it is a dominant with two of its notes delayed, and analysing it as a genuine tonic arrival will make nonsense of the phrase.

Example. A phrase in G major ends with the chords Em, C, D, G. Give the roman numerals and name the cadence.

G major is G A B C D E F#. E minor is the triad on degree 6, so vi; C is IV; D is V; G is I. The numerals are vi, IV, V, I, and the last two make a perfect cadence.

Now you. A phrase in F major ends Bb, C, Dm. Give the numerals and name the cadence.

Answer

F major is F G A Bb C D E. B flat is IV, C is V, D minor is vi. The numerals are IV, V, vi, and V going to vi is an interrupted cadence.

Voice leading: the lines matter more than the chords

Write harmony out in four parts, the soprano, alto, tenor and bass of a chorale, and a second set of constraints appears. A chord progression is simultaneously four melodies, and they have to be singable and distinguishable.

Three principles cover most of it. Keep common tones: if two consecutive chords share a note, leave it in the same voice. Move everything else the shortest distance available, which usually means by step. And resolve the tendency tones: the leading note goes up to the tonic, and the seventh of a seventh chord goes down by step, both because that is what the tritone argument above requires.

Apply all three to G7 going to C in four parts, with the bass making the cadential leap from G to C. B is the leading note and rises to C. F is the seventh and falls to E. D, having no obligation, takes the nearest note, C. The tonic chord that arrives is C, C, C, E: root tripled, third present, fifth missing altogether. That result surprises everyone the first time and is the standard textbook outcome, and the way round it, if a complete tonic is wanted, is to write the V7 incomplete instead, doubling its root and omitting its fifth so that D is not there to be resolved.

The famous prohibitions follow from the same concern. Parallel fifths and parallel octaves, two voices moving in the same direction while staying a perfect fifth or octave apart, are avoided because perfect intervals blend so completely that the two voices stop being heard as two. It is not a moral rule and not an acoustic law; it is a consequence of wanting four independent lines, and music wanting a different texture, from medieval organum to nearly all rock guitar, uses parallel fifths deliberately.

Progressions people actually use

A handful of patterns account for an enormous amount of music, and they are all instances of the ideas above.

The falling-fifth chain, usually shortened to vi, ii, V, I, is the backbone of Baroque sequences and of jazz standards, where ii7 to V7 to Imaj7 is so standard it is written "two five one" and treated as a single unit. The four-chord pattern I, V, vi, IV underlies a long list of pop songs and works because it never resolves harder than it has to: the vi where I was expected is an interrupted cadence used as a hinge rather than as a surprise.

The twelve-bar blues deserves its own paragraph because it breaks the rules productively. Its skeleton in C is four bars of C, two of F, two of C, then one of G, one of F, and two of C, with the last often turning back to G. Two things about it are irregular. The move from G to F in bars 9 and 10 is dominant to predominant, the retrogression classical style avoids. And all three chords are routinely played as dominant sevenths, C7, F7 and G7, although only one dominant seventh exists in any major key: the blues treats the sonority as a colour rather than a function, so strict functional rules will simply mark it wrong.

Leaving the key

Two devices extend the system without abandoning it.

A secondary dominant borrows the dominant of a chord other than the tonic. In C major, to arrive at G with the same inevitability that V7 gives to C, precede it with the dominant seventh of G, which is D, F#, A, C. That chord is not in C major, since F sharp is not in the scale, and it is written V7/V, "five seven of five". The F sharp is the leading note of G, and its presence tells the ear that G is momentarily behaving as a tonic. This is tonicisation: a brief nod at another key, over in a chord or two, and it works on every degree except the diminished one.

Modulation is the same idea taken seriously enough to change key. The smoothest method is the pivot chord, a chord belonging to both keys, played first as a member of the old key and then reinterpreted as a member of the new one, after which a cadence in the new key confirms the change. Moving from C major to G major, the chord A minor is vi in C and ii in G, so a progression can arrive at A minor as vi, continue to D7 and G, and by the time the cadence lands the ear has accepted G as home without noticing a seam.

Example. Modulate from C major to G major. Find a pivot chord and write a progression that establishes the new key.

C major and G major share six notes, so several chords qualify. Take A minor: vi in C, ii in G. A progression of C, F, Am, D7, G reads as I, IV, vi in C, then ii, V7, I in G, with A minor doing both jobs. The D7 is the confirmation: it contains F sharp, which C major does not have, and after it the key is no longer arguable.

Now you. Find a pivot chord for a modulation from F major to C major, and name what it is in each key.

Answer

D minor works: it is vi in F major and ii in C major. Follow it with G7 and C, which are V7 and I in the new key, and the F natural in G7 is common to both keys while the B natural in it belongs only to C major.

What functional harmony does not do

Be clear about the status of everything above. It describes the common practice of European music roughly between 1600 and 1900, generalised largely from Bach chorales because those are short, numerous and consistent. It predicts what such music usually does. It composes nothing, explains no piece's quality, and does not travel outside its repertoire.

Outside it, plenty of good music disagrees. Modal folk and rock progressions use bVII and bIII, which functional analysis labels as borrowings but which are simply the notes of the mode being used. Debussy writes parallel chords whose purpose is colour rather than function. And a great deal of music since 1900 has no tonic at all, at which point the vocabulary of this lesson has nothing to say.

What it does give is a reliable account of why a chord progression in a huge and familiar body of music pulls the way it does. The last thing standing between that and reading an actual score is everything on the page that is not a chord tone, which is the final lesson.

Reading a score

A real score contains notes that belong to no chord and marks that are not notes at all, and this lesson closes the gap between the theory built so far and a printed page.

Notes that belong to no chord

Take a bar whose harmony is plainly C major and whose melody runs C, D, E. The D is in no C major triad, and yet nothing sounds wrong. It is a non-chord tone, and once you can name the six common types you can strip a melody down to its harmony in a single pass.

A passing tone fills the gap between two chord tones a third apart, moving by step in one direction and continuing in the same direction: C, D, E, with D the passing tone. It normally falls on a weak beat.

A neighbour tone steps away from a chord tone and back to it: C, D, C is an upper neighbour, C, B, C a lower one. Also usually weak.

A suspension is the one that lands on a strong beat, and it has three stages. A note is sounded as part of one chord (preparation), held while the harmony changes underneath so that it becomes a dissonance (suspension), then falls by step into the new chord (resolution). Suspensions are named by the intervals above the bass, so the commonest is the 4 to 3. They are the main way European music creates a dissonance where it wants to be noticed.

An anticipation is the mirror image: a note of the next chord arrives early, over the old harmony, usually just before a cadence. An appoggiatura is leapt into and resolves by step, on a strong beat, so it sounds like a suspension that arrived without preparation. An escape tone does the opposite, stepping away from a chord tone and then leaping.

The classification is entirely a matter of how the note is approached and how it is left, so the procedure is mechanical: look at the note before and the note after, and check whether each move is a step or a leap and whether the note is metrically strong or weak.

Example. A bar is harmonised with F major (F A C). The melody plays A, G, F, E, F in eighth notes, starting on beat 1. Which notes are non-chord tones, and of what kind?

The chord tones are F, A and C. So G and E are the outsiders. G, on the second eighth, is approached by step down from A and left by step down to F, so it is a passing tone. E, on the fourth eighth, steps down from F and steps back up to F, so it is a lower neighbour tone. Both fall on weak eighths, which is what the labels predict.

Now you. A bar is harmonised with G major (G B D) over a G bass. The melody holds a C on beat 1, tied over from the previous bar, then moves to B on beat 2 and D on beat 3. Name the non-chord tone and its type.

Answer

C is not in G B D. It was sounded under the previous harmony, held into this one where it is dissonant, and resolves down by step to B, so it is a suspension. Since C is a fourth above a G bass resolving to a third, it is a 4 to 3 suspension.

Loudness and attack

Marks for loudness and for attack control what the notes sound like, and neither family is precise.

Dynamics are relative. The letters run pp, p, mp, mf, f, ff, from soft to loud, with ppp and fff at the extremes, and they mean nothing in decibels: an f on a solo flute and an f in an orchestral tutti are worlds apart. Gradual change is written cresc. and dim., or drawn as opening and closing hairpins. Sudden marks include sf for a single stabbed note and fp, loud then immediately soft.

Articulation says how each note begins and ends. A dot above the note is staccato, detached, roughly half its written length. A short horizontal line is tenuto, held for full value and slightly leaned on. An arrowhead is an accent, a wedge is marcato. A slur asks for its notes to be joined, which on a wind instrument literally means one breath and no re-tonguing. A fermata, the eye-shaped mark, means hold the note longer than written, and how much longer is the conductor's decision.

Tempo words

Tempo words are older than metronome marks and vaguer. In rough order of speed: largo around 40 to 60 beats per minute, adagio 66 to 76, andante 76 to 108 (it means "walking"), moderato 108 to 120, allegro 120 to 168, presto 168 to 200. Treat those numbers as a modern convention rather than a fact, since editions disagree with each other and with Maelzel's original table, and since the words were understood as characters before they were understood as speeds: allegro means cheerful, not fast. Changes are marked rit. or rall. to slow down, accel. to speed up, a tempo to return.

The shape of a score, and the parts that lie

An orchestral score stacks its staves in a fixed order, top to bottom: woodwind, then brass, then percussion, then any keyboard or voices, then strings. Within each family the highest instrument is on top, and the order never changes, so a conductor knows where to look without reading the labels.

The trap is that many parts are written in a different key from the one they sound. A transposing instrument is one whose written C produces some other pitch, and the reason is fingering: a clarinettist who learns one set of fingerings can then play a B flat, an A or an E flat clarinet from parts that all look the same. The commonest are the B flat instruments, clarinet and trumpet, which sound a major second below what is written; the horn in F, a perfect fifth below; the alto saxophone in E flat, a major sixth below; and the tenor saxophone in B flat, a major ninth below. Piccolo sounds an octave above and double bass an octave below, which are transpositions too, if trivial ones.

Two rules keep this straight. To go from written to sounding, transpose down by the instrument's interval. To go from sounding to written, transpose up. And a transposing part carries its own key signature: with the orchestra in C major, a B flat clarinet part is written in D major, two sharps, because everything it reads is a tone higher than it sounds.

Example. A B flat trumpet plays a written D5. What pitch sounds, and what would the trumpet have to read to sound a concert F4?

Written to sounding is down a major second, so a written D5 sounds C5. Sounding to written is up a major second, so to produce a concert F4 the player must read G4.

Now you. A horn in F reads a written C5. What sounds? And what must an alto saxophone read to sound a concert E flat 4?

Answer

The horn sounds a perfect fifth lower than written, so written C5 sounds F4. The alto saxophone sounds a major sixth lower, so to sound E flat 4 it must read a major sixth higher, which is C5.

Eight bars, analysed from nothing

Here is a chorale-style phrase in G major, 4/4, its melody four quarter notes to a bar, its bass one or two notes, and its inner voices left as chord names. Nothing has been analysed yet.

BarMelodyBassChords
1D5 E5 D5 B4G2G
2C5 B4 A4 F#4C3, D3C, D
3G4 B4 C5 E5E3, A2Em, Am
4F#5 E5 D5 D5D3D
5D5 B4 E5 C#5G2, A2G, A7
6D5 D5 C5 A4D3D, D7
7B4 B4 C5 C5E3, A2Em, Am
8B4 B4 A4 G4D3, G2G/D, D7, G

Start with the key. The signature would be one sharp, and the piece begins and ends on G with F sharps throughout, so it is G major rather than E minor. The diatonic triads are therefore G, Am, Bm, C, D, Em and F# diminished, numbered I, ii, iii, IV, V, vi and vii°.

Now the numerals, chord by chord. Bar 1 is G, so I. Bar 2 is C then D, so IV then V. Bar 3 is Em then Am, so vi then ii. Bar 4 is D, so V, and it lasts a whole bar and ends the phrase: that is a half cadence, and it means bars 1 to 4 form an antecedent, a musical question.

Bar 5 begins again on I, then reaches A7, which is not in G major at all, since it contains C sharp. Its root A is degree 5 of D and the C sharp is D's leading note, so it is the dominant seventh of the dominant, V7/V, tonicising the D that arrives in bar 6. Bar 6 is V, turning into V7 in its second half by adding C. Bar 7 does not resolve that V7 to I: it goes to Em, vi, so bars 6 to 7 make an interrupted cadence. The rest of bar 7 is Am, ii.

Bar 8 is the real ending. G/D is a G chord with D in the bass, second inversion, so I6/4, sitting on a strong beat above the dominant bass note. It resolves into D7, V7, and then to G, I: a cadential six-four followed by a perfect cadence. The six-four is not really a tonic chord, since B and G above the D bass are a sixth and a fourth resolving down to A and F sharp, a fifth and a third, which is a pair of suspensions over the dominant.

Finally the melody, where almost every note is a chord tone and the exceptions are the interesting ones. The E5 on beat 2 of bar 1 steps up from D5 and back to it, an upper neighbour. The B4 on beat 2 of bar 2, over a C chord, is a passing tone between C5 and A4, and the E5 on beat 2 of bar 4 passes between F sharp 5 and D5. The C sharp 5 ending bar 5 is a chord tone of A7 behaving as a leading note, and it rises to D5. The C5 on beat 3 of bar 6 is the seventh of D7 and duly falls to B4 at the start of bar 7.

The whole thing is a period: two four-bar phrases, the first ending in a question and the second answering it, which is the commonest eight-bar shape in the repertoire. Written out compactly, the analysis is: I, IV, V, vi, ii, V, I, V7/V, V, V7, vi, ii, I6/4, V7, I.

Example. In that excerpt, why is the A7 in bar 5 written with a C sharp rather than a D flat?

Because spelling records function. The note is the leading note of D and rises a semitone to it, so it must be written as a raised C, and a D flat would say that the note falls to C. The two sound identical in equal temperament and mean opposite things on the page.

Now you. Suppose bar 8 were changed so that the last chord were Em instead of G. What would the final cadence be called, and what would the melody's last note, G4, be doing?

Answer

V7 to vi is an interrupted, or deceptive, cadence, so the phrase would refuse to close. G4 is a chord tone of E minor, being its third, so the melody would still fit, but a phrase ending on degree 1 harmonised as the third of vi sounds unfinished, which is the whole point of the device.

What the analysis will not tell you

Everything above is worth doing and none of it is a theory of music. Four limits are worth carrying away.

Roman-numeral analysis is descriptive. It reports what a passage does in the vocabulary of one style. It does not say why the piece is good, does not generate a next chord, and applied to the wrong repertoire it produces labels that are technically correct and musically empty.

Analysis is not unique. Choosing which notes are structural and which are decoration is a judgement, and competent analysts disagree, most often about whether a chord is a real harmony or a by-product of passing notes. Bar 8's six-four is a mild case: it can be read as a tonic chord in second inversion or as a decorated dominant, and this lesson picked the second reading because of how it resolves.

The theory was built from a narrow sample. Its rules are generalisations from European music of roughly 1600 to 1900, and even inside that window composers break them knowingly. Outside it, the machinery gets steadily less useful: a little for jazz, less for folk and rock, almost nothing for music not organised around a tonic.

Finally, notation is a set of instructions, not a recording. Tempo, dynamics and articulation are approximate, rhythm is bent in performance, and the difference between two good performances of one page is larger than anything the page specifies.

What you can now do is open a score you have never seen, work out its key and metre, name its intervals and chords, put roman numerals under a phrase, find the cadences, and say why the harmony pulls where it does. The rest is repertoire: take eight bars of something you like and do to it exactly what was done above.

Music Theory, from libre.university