Sign in

Libre University uses your GitHub account. Signing in is only needed to sit a final test, so the score is kept on your profile.

Reading kanji at scale

Seven lessons of grammar have left you able to build a sentence more complicated than most of what appears in a newspaper, and still unable to read one, because the obstacle has stopped being grammar. This lesson is about the writing system, and specifically about the two things that let a reader make progress against thousands of characters rather than memorising them one at a time.

What the list is, and how it got that size

Japanese has no fixed inventory of characters, but it does have an official recommended set, and knowing its history explains a lot about what you will and will not meet.

In 1946 the occupation-era government published the 当用漢字とうようかんじ, 1850 characters, explicitly as a restriction: anything outside it was to be avoided or rewritten in kana. That list is why some compounds are still written with a mismatched character or spelled half in kana. In 1981 it was replaced by the 常用漢字じょうようかんじ, 1945 characters, and the framing changed from a limit to a guide for official documents, newspapers and broadcasting. On 30 November 2010 the list was revised again: 196 characters were added and 5 removed, giving 1945 plus 196 minus 5, which is 2136, the figure in use today. Most of the additions were characters that word processors had made easy to write and that people were using anyway, such as うつ and だれ.

Inside that set sits the 教育漢字きょういくかんじ, the characters taught in the six years of primary school. It stood at 1006 for decades and was raised to 1026 in a revision applied from 2020, the twenty additions being characters used in prefecture names. So a Japanese child finishes primary school knowing 1026 characters, roughly 48 per cent of the jōyō set, and finishes lower secondary school with the rest.

Two honest qualifications. First, the jōyō list is a guide and not a boundary. Personal names may use roughly 860 further characters from a separate permitted list, and literature, menus and shop signs cheerfully use characters from neither. Second, and more usefully, knowing a character is not the same as knowing a word. The jōyō set is a list of building blocks, and the thing you actually read is vocabulary.

The phonetic half of a character

The single most useful fact about kanji is that the majority of them are not pictures. They are compounds of two parts: one indicating a meaning domain and one indicating a sound.

Xu Shen said so in the year 100 CE. His 説文解字せつもんかいじ analysed 9353 characters and classified 7697 of them, about 82 per cent, as phono-semantic compounds, where one component gives the sense and the other the pronunciation. That proportion has only risen since, because new characters have almost always been coined the same way.

The clearest series to see it in is はん, whose on reading is ハン.

CharacterComponentsOn readingA word
ばんwood + はんハン, バン黒板こくばん, blackboard
はんslice + はんハン出版しゅっぱん, publishing
はんshell, money + はんハン販売はんばい, sales
はんfood + はんハンはん, cooked rice
はんearth + はんハン急坂きゅうはん, steep slope
はんhill + はんハン大阪おおさか, Osaka

Six characters, one sound, and in every case the left-hand part tells you roughly what field the word belongs to: wood, paper, money, food, land. The せい series works the same way and is even more common: せい, せい, せい and せい all carry セイ from せい, and their left-hand parts give water, sun, rice and words respectively. じょう is in the series too, with the heart radical on the left, and reads ジョウ, which is the voiced variant of the same borrowed syllable rather than a different sound. せい shows the other thing to watch for: it is built from the same せい, but here the phonetic sits on the left and そう on the right, so the useful habit is to look for a component you know anywhere in the character rather than only on the right. きよい, clear, is water plus the sound セイ, and its on reading is セイ, exactly as advertised.

This is worth internalising because it changes what learning a new character costs. If you know はん and you meet はん for the first time, you do not have to learn a reading; you have to check a guess.

Example. You meet こう for the first time. Its left part is the thread radical and its right part is こう, whose on reading is コウ. What do you guess, and how do you check?

Guess コウ for the reading and something to do with thread, cloth or colour for the meaning. Both are right: こう means crimson, and 紅茶こうちゃ, literally red tea, is the ordinary word for black tea. The check is to look for a compound you already know that contains it, which is the fastest confirmation available and also the only one that gives you a word rather than a character. Note that the same phonetic こう also sits in こう, こう and こう, all コウ, so the series is a strong one.

Now you. Given that has the on reading ジ, predict the on readings of , , たい and とく, then say what that tells you about the method.

Answer

You would predict ジ for all four. and are ジ and the prediction holds; たい is タイ and とく is トク, and it fails. The series is one of the weak ones, and it is a good demonstration of what phonetic components are: a strong prior, not a rule. Roughly speaking, a phonetic hint is worth taking when you have already seen two or three characters in the series agree, and worth checking always. What the failures do not do is make the method worthless, because a guess you verify in five seconds is still much cheaper than rote memorisation.

On and kun, and what okurigana tells you

The beginner course introduced the two kinds of reading: on readings, which are Japanese approximations of the Chinese pronunciations the characters arrived with, and kun readings, which are native Japanese words the characters were attached to. The question a reader faces is which one a given occurrence is using, and there are two reliable signals.

The first is okurigana, the kana trailing the character. Native Japanese words inflect, and the inflection has to be written in kana, so trailing kana are the fingerprint of a kun reading. べる has okurigana and is read た; 食事しょくじ has none and is read しょく. あたらしい is あたら, 新聞しんぶん is しん. たかい is たか, 高校こうこう is こう. This one test resolves the majority of cases you will meet in ordinary text.

The second is company. A character standing next to another kanji, with no kana between them, is usually taking its on reading, because two-character compounds were mostly borrowed as units from Chinese. 電車でんしゃ, 銀行ぎんこう, 勉強べんきょう, 試験しけん: all on plus on.

The exceptions to the second rule are common enough to have their own names, which tells you they are not rare. A compound reading on then kun is called 重箱読じゅうばこよみ after the word 重箱じゅうばこ itself, a stacked lunch box, read じゅう plus はこ. Reading kun then on is 湯桶読ゆとうよみ after 湯桶ゆとう, and it covers a set of extremely common words: 場所ばしょ, 手本てほん, 見本みほん, 台所だいどころ. There is no rule that predicts which compounds are mixed. They have to be learned as words, which is the recurring moral of this lesson.

Example. You meet 下手へた in text, having previously learned した and . Why does the reading fail to be predictable, and what should you conclude?

Neither part is using an obvious reading of the sort the rules give: 下手へた is not したて, and it is not the on-on かしゅ, and its meaning, unskilled, is not the sum of down and hand. It is a fixed word whose spelling was assigned to it, and the characters are a label rather than a derivation. The conclusion is the practical one that runs through this whole lesson: characters are for guessing at words you half know, and words are what you actually learn. A dictionary lookup here is not a failure of method.

Now you. Predict the readings of 花見はなみ and 見学けんがく from the rules above, and say which rule each one follows.

Answer

花見はなみ, flower viewing, is kun plus kun, はな and み, and is a compound of two native words rather than a borrowing. 見学けんがく, a study visit, is on plus on, けん and がく, and follows the ordinary compound rule. The character is doing both jobs, and there is no way to tell from the character alone which is which; what tells you is that 花見はなみ is a native cultural word and 見学けんがく is a formal Sino-Japanese one. Register is a surprisingly good predictor: the more official the word feels, the more likely it is on-on.

What 2136 characters actually buys

Frequency counts of Japanese newspaper text, made repeatedly since the 1960s, differ in detail and agree on the shape. The commonest 500 characters account for roughly 80 per cent of all kanji occurrences; the commonest 1000 for about 95 per cent; and 2000 for a little over 99 per cent. In other words a little under half the jōyō set, 1000 characters out of 2136, or 47 per cent, does 95 per cent of the work.

That number is encouraging until you do the arithmetic on what it means for reading. Coverage is measured in occurrences, not in words, and most words contain two characters. If your knowledge is 95 per cent complete and the gaps fall independently, the chance of knowing both characters of a two-kanji word is 0.95 squared, which is 90 per cent. At 99 per cent coverage it is 98 per cent. And in an article of 1200 characters, 5 per cent unknown is 60 characters you cannot read, or roughly one every two lines.

This is why learners at exactly this stage so often feel that their kanji study is not paying. It is paying; the arithmetic of composition simply eats a lot of it. The way out is not more characters but more words, because the gaps that matter are lexical.

Example. You know 1000 characters, giving about 95 per cent coverage. A page has 800 characters of running text, of which about 40 per cent are kanji. How many unknown characters should you expect, and how many two-kanji words will you fail on?

The kanji on the page number 800 times 0.4, which is 320. Five per cent of those are unknown, giving 16 unknown characters on the page. If those 320 characters make up about 160 two-character words, the proportion you know completely is 0.95 squared, or 90.25 per cent, so you should expect to fail on about 16 words. Sixteen lookups per page is slow but survivable, which is roughly the experience of a reader at this level, and it improves quickly because the unknown characters are not randomly distributed: the ones you meet are the ones that recur.

Now you. At 99 per cent coverage, what happens to those two figures on the same page?

Answer

Unknown characters fall from 16 to 320 times 0.01, which is about 3. Two-kanji words known completely rise to 0.99 squared, or 98.01 per cent, so failures fall from about 16 words to about 3. Going from 95 to 99 per cent coverage means learning roughly another thousand characters, and it cuts the lookups by a factor of five. That is the honest shape of the last stretch: a large amount of work for a change that is nonetheless the difference between decoding a page and reading it.

Reading a character you have never seen

Put the pieces together into a procedure, because having one is what separates a reader who makes progress from one who stops at every unknown character.

Look first at whether kana follow it. If they do, the reading is native, the word is probably one you know spoken, and the okurigana often identifies it outright: a character followed by べる in a sentence about a meal is べる whether or not you recognise the character.

If it sits against another kanji, assume on readings for both and look for a component you know on the right or the bottom. That component is the sound more often than not, and one candidate reading plus the surrounding sentence is usually enough to recover a word you have heard.

Then check, and check by looking up the word, not the character. A character learned alone is almost useless, since you cannot say when it takes which reading; the same character inside three words you can use is knowledge you keep.

Finally, accept a miss. Native readers do not know every character either, which is why furigana exists at all, why newspapers gloss rare characters, and why the 2010 revision was partly an admission that people were using characters nobody could write by hand.

Vocabulary

WordReadingMeaning
常用漢字じょうようかんじjouyou kanjithe 2136 recommended characters
出版しゅっぱんshuppanpublishing
販売はんばいhanbaisales
黒板こくばんkokubanblackboard
紅茶こうちゃkouchablack tea
場所ばしょbashoplace
台所だいどころdaidokorokitchen
見本みほんmihona sample
花見はなみhanamiblossom viewing
見学けんがくkengakua study visit
下手へたhetaunskilled
意味いみimimeaning
辞書じしょjishodictionary
調しらべるshiraberulook up, investigate

That is the last lesson away from grammar. The rest of the course goes back to the verb and to the social information Japanese packs into it, starting with something English leaves entirely to context: who did whom a favour.