How to study Japanese vocabulary: seven methods compared

Every learner eventually discovers that Japanese grammar is finite and Japanese vocabulary is not. You can finish a grammar syllabus; you cannot finish words. That makes vocabulary method the highest-leverage decision in your study plan — and the one most people never make deliberately, drifting instead into whatever their first textbook happened to do. This guide compares seven methods honestly, including what each is bad at, then covers the principles that decide what actually sticks, what to do about the words that refuse to stick, and the real arithmetic of getting to 2,000, 6,000 or 10,000 words.

1. The seven methods at a glance

"Retention" below means retention per hour invested, not retention in the abstract — a method that gets a word to stick after forty minutes of work is not automatically better than one that half-sticks it in two. Every method here works for someone; none works for everything.

MethodTime cost per wordRetention per hourBest stageMain failure mode
Bulk wordlistsVery low (5–10 sec)Low aloneFirst exposure; pre-teaching before a lesson; exam cramYou learn the list order, not the words
Flashcards + SRSModerate (30–60 sec to set up, ~8 sec per review)HighEvery stage; the default engineLeeches and deck bloat; knowing the card rather than the word
Sentence miningHigh (2–4 min per item)Very highN4 upward, once you can read something realCard-making becomes the hobby and eats the study time
Reading-driven acquisitionLow per encounter, but needs 8–20 encountersLow early, very high laterN3 upward, or anyone past ~1,500 wordsBelow the comprehension threshold it teaches almost nothing
Listening immersionNear zero marginal (commute, chores)Poor for new words, high for speedAny stage, as a consolidatorAudio you don't understand is background noise, not input
Writing / output practiceHighest (minutes per word, in use)Highest for productionAny stage, in small dosesWithout correction you fossilise your own errors
Mnemonics / keyword methodHigh (30–90 sec to invent one)Very high, for the specific wordTargeted use on hard words and kanji readingsApplied to everything, you end up recalling mnemonics, not Japanese

2. What each method is actually for

Bulk wordlists have a bad reputation they only half deserve. Reading straight down a list of 50 words is genuinely poor as a way to learn them, but it's excellent as a way to meet them. Twenty minutes with the N5 list before you start drilling means every card afterwards is a second encounter rather than a first, and second encounters are markedly cheaper. The catastrophic version is treating the list as the whole method: because the words always appear in the same order, you build order-dependent recall and then discover that 出発 (しゅっぱつ, departure) is only retrievable when it follows 到着 (とうちゃく, arrival).

Flashcards with spaced repetition are the workhorse, and for good reason: they are the only method that systematically schedules recall attempts at the point where memory is fading, which is precisely where the strengthening happens. The mechanics and the scheduling maths are covered in detail in the spaced repetition guide; what matters for method choice is the trade-off. SRS is unbeatable at getting a large number of form–meaning links into memory cheaply, and comparatively weak at teaching you how a word behaves in a sentence. A card that says 遠慮 → "reserve, restraint" will let you recognise 遠慮 (えんりょ) on a reading test and will not help you produce 遠慮なくどうぞ ("please, don't hold back"). Use it for coverage, not for depth.

Sentence mining is the depth method. Instead of harvesting words, you harvest sentences you met in real material and understood except for one piece. Rather than a card for 割引 (わりびき, discount), you keep 今日は全品20%割引です (きょうは ぜんぴん にじゅっパーセント わりびきです, "everything is 20% off today"), which carries the word, its typical company, and its register in one unit. Retention is excellent because the sentence supplies retrieval cues that a bare gloss cannot. The cost is real: two to four minutes per item once you include finding it, checking the reading, and typing it up. Most people who abandon sentence mining do so because they tried to mine everything.

Reading-driven acquisition is how you eventually learn most of your words, and it's almost useless before a threshold. Research on second-language reading consistently finds that incidental acquisition needs many encounters — commonly cited figures run from about eight to twenty exposures before a word is reliably known — and that comfortable independent reading requires knowing roughly 98% of the running words, which is about two unknowns per hundred. Below that, you're decoding rather than reading, and almost nothing sticks. This is the honest reason "just read more" fails as beginner advice and becomes excellent advice around N3.

Listening immersion is widely oversold and genuinely valuable, but for a different job than people think. It rarely teaches new words. What it does superbly is turn slow knowledge into fast knowledge: a word you can recall in four seconds is useless in conversation, and repeated listening is what compresses that to a quarter of a second. It also fixes the pronunciation of words you learned by eye — an enormous problem in Japanese, where you can carry a wrong reading for 一日 or 十分 for years. Passive audio you don't understand does approximately nothing; comprehensible audio does a great deal.

Output practice — writing sentences, journalling, speaking — is expensive per word and irreplaceable, because it is the only method that surfaces the gap between what you recognise and what you can use. You will not discover that you can't reliably choose between 借りる (かりる, to borrow) and 貸す (かす, to lend) by reviewing flashcards; you'll discover it the first time you need to say one of them. Five sentences a week using words from your current batch is enough to expose most of these gaps.

Mnemonics are a scalpel, not a hammer. Inventing an image for every word is unsustainable and self-defeating, because a mnemonic is an extra retrieval step you eventually want to discard. But for a word that has failed you ten times, thirty seconds spent building a vivid, specific, slightly absurd link is the single most effective intervention available. Kanji readings are the other legitimate target, especially where a phonetic component is doing predictable work and the mnemonic only has to cover the exception.

3. The principles that actually drive retention

Methods are surface. Underneath them sit a handful of principles, and a mediocre method applied with the principles beats a fashionable method applied without them.

Interleaving deserves a Japanese-specific note, because the language hands you ready-made interleaving material: transitive/intransitive verb pairs. 開ける (あける, to open something) and 開く (あく, for something to open); 出す (だす, to put out) and 出る (でる, to go out); 始める (はじめる, to begin something) and 始まる (はじまる, for something to begin). Studied in separate blocks these merge into mush. Studied together, with the pair explicitly contrasted and a particle in the prompt — ドアを開ける versus ドアが開く — they reinforce each other.

4. Leeches: the words that refuse to stick

In any deck of a thousand words, a small minority — often around five per cent — will consume a wildly disproportionate share of your review time. These are leeches, and most SRS tools flag them automatically at around eight failures. Ignoring them is expensive: fifty leeches at ten extra reviews each is five hundred reviews you didn't need to do. Treating them is quick, but only if you diagnose first, because leeches have distinct causes.

  1. Interference — the most common cause. The word isn't hard; it's colliding with a similar one. 訪ねる (たずねる, to visit) and 尋ねる (たずねる, to ask) share a reading. 貸す and 借りる are logical opposites you can name but can't produce under pressure. The fix is counter-intuitive: stop trying to separate them and put both on one card, with the contrast stated. You cannot un-link things that are already linked, so link them deliberately and correctly.
  2. A bad prompt. A card reading 適当 → "suitable" will fail forever, because 適当 (てきとう) also carries the everyday sense of "slapdash, done half-heartedly", and your brain is refusing a gloss it knows is wrong. Rewrite the card with a sentence that pins one sense down.
  3. No hook. Abstract words with no image and no personal relevance have nothing to attach to. This is where mnemonics earn their cost, and where a personal example sentence — one about your own job, city or family — outperforms a textbook one.
  4. It's genuinely low value. Sometimes a word has failed fifteen times because you never encounter it. Delete it. Deleting a card is not failure; it's reallocating twenty future reviews to words you'll actually meet. You'll pick it up from reading if it matters.

Set a fixed appointment for this — fifteen minutes a week to look at whatever your review queue has flagged and apply one of the four fixes to each. It's the highest-return quarter-hour in the whole system, and almost nobody does it.

5. A week that combines the methods

No single method covers coverage, depth, speed and production. A workable week uses three or four in fixed slots, so that the choice is already made when you sit down. Here is a roughly 35–45 minute-a-day structure for an intermediate learner; scale the numbers, keep the shape.

DayWhat you doTime
Mon / Wed / FriFull session: clear the review queue, then add 10–15 new words. Finish with 10 minutes of reading at a level where you meet 2–5 unknown words per page40 min
Tue / ThuReviews only, no new words — this is what stops the queue compounding. Add 20 minutes of listening you genuinely follow, on the commute if that's what you have20 min + commute
SaturdayDepth session: mine 5–8 sentences from the week's reading or viewing, then write five sentences of your own using words from this week's batch45 min
SundayReviews, then the 15-minute leech audit. Take the rest of the day off — the schedule needs slack or the first missed day becomes a missed month25 min

Two details make or break this. First, the no-new-words days are not a concession, they're load management: new words are what generate future reviews, so two dry days a week keeps the queue from growing faster than you can clear it. Second, everything is anchored to an existing habit — the commute, the coffee, the train. A daily session that runs at a fixed moment survives; one that runs "when I have time" does not.

6. The honest arithmetic: how long to 2,000, 6,000, 10,000

Vocabulary targets are usually quoted without the review cost attached, which makes them look far cheaper than they are. Here is the missing half. In a typical spaced-repetition schedule, each new word generates roughly eight to twelve reviews on its way to being securely known. That means a steady intake of N new words a day settles at somewhere around five to ten times N reviews a day. Ten new words is not ten items of work; it's ten items plus fifty to a hundred reviews.

TargetNew words/dayStudy days neededDaily time at steady stateRealistic elapsed time
2,000 (solid beginner, N4-ish range)1020025–35 min9–12 months at six days a week
6,000 (comfortable intermediate, N2-ish range)1540035–50 min2–3 years cumulative from zero
10,000 (advanced, N1-ish range)2050050–75 min3–5 years cumulative, and few people sustain 20/day throughout

The "study days" column is arithmetic; the "elapsed time" column is arithmetic plus honesty. Nobody studies 365 days a year, retention is never 100%, and illness, work and holidays remove weeks at a time. A 15–25% haircut on the perfect-execution number is the normal experience, not a sign that something has gone wrong.

Two caveats on the targets themselves. First, the JLPT has not published an official vocabulary list since the test was restructured — the old 出題基準 gave roughly 800 words for the lowest level and about 10,000 for the highest, and every modern figure you see is an estimate derived from past papers, so treat level-to-word-count mappings as approximate. Second, "knowing 6,000 words" hides an enormous range: 6,000 words at recognition level is a different achievement from 6,000 you can produce. For planning purposes, assume your production vocabulary is a third to a half of your recognition vocabulary unless you've trained the other direction deliberately.

On this site the material maps roughly onto those tiers: the basic module holds 1,500 everyday words, the JLPT module 2,904 across the five levels, and the JPT module 5,862 weighted towards business and workplace language — a little over 10,000 in total, which is enough to run the whole arc above without changing sources. The study page shows how the sets break down if you want to plan a route through them.

7. Measuring progress without fooling yourself

The metric almost everyone uses — words added — is the one metric that measures nothing. Adding is free; retaining is the work. Four better measures:

Whatever you measure, measure it the same way each time and write it down. Memory of your own progress is unreliable in both directions: on a bad week you'll believe you've learned nothing, and after a good session you'll believe you've learned a hundred words. A three-line log beats both impressions. You can run the recall drills themselves from the study page, and push anything the log throws up back into your review deck.

8. Choosing, in one paragraph

If you're below about 1,000 words, weight your time towards lists and spaced repetition and don't feel guilty about isolation learning — you need coverage before context becomes affordable, and reading is not yet giving you a return. Between roughly 1,000 and 3,000, keep the SRS engine running but start shifting the balance towards sentences and reading, and add the production direction for your most useful words. Above 3,000, invert the ratio: reading and listening become your main acquisition channels, with the SRS demoted to a net that catches what reading alone won't secure. The mnemonic scalpel and the weekly leech audit stay constant throughout. The one thing that stays constant regardless of level is the scheduling — an hour on Sunday will never beat twenty minutes on six days, no matter which method fills the twenty minutes.

Pick two methods, not seven: one engine for coverage and one for depth. Run the engine daily, run depth weekly, audit your leeches every Sunday, and measure mature words rather than added words. Do that for a year at ten to fifteen new words a day and you'll be somewhere past 2,500 words — which is not a heroic number, and is enough to change what Japanese feels like.