The idea of "learn a hundred characters and start reading" is one of the most persistent in the Chinese learning community. It appears in course titles, app promises, and forum advice. And it's almost true—if by "reading" you mean recognizing symbols on a page, not understanding text. Let's break down why the first 100 characters are a real milestone, but not the one usually being sold, and how to choose and order them to actually get closer to living Chinese text.
The arithmetic you should start with
According to corpus linguistics data on modern Chinese (specifically, Jun Da's frequency lists based on the Modern Chinese Character Frequency List corpus), the first 100 most frequent characters cover about 42% of all tokens in written texts. Sounds impressive: almost every other character on the page is familiar to you. But here's the trap: the familiar ones are precisely the most frequent—that is, functional and grammatical characters—particles, pronouns, connectives, basic verbs. And the incomprehensible 58% are exactly the nouns, adjectives, and specialized verbs that carry the main meaning.
Imagine a Russian text where you understand all the "and," "in," "not," "I," "was," "this," but don't understand any of the nouns. Would you get the meaning? No. That's exactly why the promise "100 characters = reading" is a methodological exaggeration. The first 500 characters cover about 75% of text, 1000 cover about 89%, and only at the 2000–3000 character mark does what can be called independent reading of non-adapted materials begin. A hundred characters isn't the finish line, but the first working zone where it makes sense to start encountering written language—not the foundation on which reading already stands.
A character is not a word, and this matters more than it seems
The main methodological error of early lists is teaching characters as lexical units. In modern Chinese, most words are disyllabic—that is, they consist of two characters. Take 学: by itself it's almost never used in living speech, but it appears in 学习 (to study), 学生 (student), 大学 (university), 学校 (school). Having learned 学 as "to study," a student encounters 学生 in text and doesn't understand whether this is a new word or two words in a row.
Hence the first practical rule: each character from the first hundred should be learned together with 2–3 high-frequency bigrams it appears in. Not 我 as "I," but 我 + 我们 (we) + 我的 (my). Not 好 as "good," but 好 + 你好 (hello) + 好的 (okay, agreement) + 好看 (beautiful). This doubles the volume of material to memorize, but exponentially increases the ability to recognize words in text—and that's precisely the goal.
Order matters more than composition
If you take any publicly available top-100 frequency list, the composition will be roughly the same: pronouns (我, 你, 他, 她), basic particles (的, 了, 是, 不), numbers (一, 二, 三, 十), temporal connectives (在, 有, 和, 就), simple verbs (来, 去, 看, 说, 想). Disagreements between sources are at the level of 5–10 characters—not fundamental. The truly debatable question is what order to take them in.
There are three main approaches. Strictly by frequency: you go top to bottom through the corpus list. Plus—maximum speed of text coverage. Minus—the first 20 characters are almost entirely functional (的, 一, 是, 不, 了, 人, 我, 在...), and you can't assemble a single meaningful sentence from them without vocabulary you don't know yet. Motivation drops in the second week.
By graphic complexity and radicals: first simple pictograms and basic radicals (人, 日, 月, 口, 山, 手), then derivatives. Plus—better systematic understanding of character structure forms. Minus—many graphically simple characters are low-frequency, and you end up knowing 山 and 水 but not recognizing the most frequent particles in text.
Hybrid, thematic-frequency: material is divided into semantic blocks (pronouns, numbers, time, basic verbs, functional particles), within each block—in descending order of frequency. After each block—immediate assembly of simple sentences. I consider this approach optimal for the first hundred characters: it provides both rapid coverage and the ability to compose real phrases from the first week.
Working structure of the first hundred
If you break down a hundred characters by functional blocks, the picture becomes clearer. Approximate distribution: 8–10 pronouns and demonstratives (我, 你, 他, 她, 它, 我们, 你们, 他们, 这, 那), 12–15 functional particles and connectives (的, 了, 是, 不, 在, 和, 也, 都, 就, 还, 把, 被), 10–12 numerals and measure words (一–十, 百, 千, 个, 些), 15–20 basic verbs (有, 来, 去, 看, 说, 想, 吃, 喝, 做, 走, 坐, 写, 读, 学, 工作), 10–15 temporal and spatial markers (今天, 明天, 昨天, 年, 月, 日, 上, 下, 里, 外, 前, 后), the rest—adjectives, question words (什么, 谁, 哪, 怎么) and several high-frequency nouns (人, 家, 学校, 中国).
Note: even within the "100 characters," part of the material consists not of individual characters but of bigrams assembled from already learned ones. 今天 = 今 + 天, 我们 = 我 + 们. This isn't cheating—this is how Chinese actually works: a hundred characters give you a vocabulary of approximately 150–200 words.
When you can open your first text
The readiness criterion isn't "I've learned 100 characters," but something stricter: I recognize each of the 100 in any graphic context (different fonts, handwriting, size), know at least one bigram for each, and can read a simple sentence made of these characters aloud without stumbling. Only after this does it make sense to go to real texts—and start not with news, but with graded readers at HSK1–2 level (for example, Mandarin Companion series, Chinese Breeze), with captions in Pleco, with short posts in educational Telegram channels, or with subtitles of children's cartoons. Trying to open People's Daily or contemporary prose with a hundred characters is a way to bury motivation in one evening.
What doesn't work, though it sounds plausible
A popular recommendation is learning characters through "picture mnemonics": 木 as a tree, 好 as a woman with a child. For the first 30–50 characters with transparent pictographic etymology, this works. But by the middle of the first hundred, mnemonics start becoming forced, and for 80% of modern characters they're simply false: most characters aren't pictograms but phonosemantic compounds, where one part gives approximate pronunciation and the other gives semantic field. Relying on visual associations as the main method at the start creates a habit you'll later have to unlearn.
What does work: spaced repetition (Anki, Pleco flashcards) with cards showing not isolated characters but characters as part of bigrams and example sentences; daily volume of 5–7 new characters, no more; regular handwriting practice for at least the first two months (even if the final goal is only reading, not writing)—motor memory stabilizes recognition.
A hundred characters isn't the key to the door of reading, but a step after which you can see how many more steps lie ahead. But if you take this step with the right order and connection to words rather than individual characters, the next five hundred come incomparably easier. On the Tomyo blog, I continue collecting notes on how Chinese learning works for adults—without popular oversimplifications and without promises that fall apart in the second month.




