Lemma vs word form: why your 2000 is not mine
Vocabulary size only makes sense if you know the unit. Three units, English and Spanish examples, and honest language-map progress for EN and ES learners.

People compare vocabulary like bank balances. "I have 2,000 words." "Research says B2 needs 3,500." "My app says 8,000." Those numbers look comparable. They almost never are.
The missing question is always the same: what counts as one word? Same number, different rulers - and ego fights start right there.
"My 2000 words" means little until you know the unit. A word form is the spelling you see (goes, went). A lemma is the dictionary headword (go). A word family bundles related derivatives (nation, national, nationalize). CEFR has no official word quota; orientation charts usually sit closer to frequent lemmas. Apps and research often use different rulers - so two honest "B2 needs 3000" claims can still disagree.
Below: the three units in plain language, English and Spanish examples, why three people with "2000 words" are not rivals, how CEFR charts and app catalogues diverge, and how personal progress on a language map works without pretending to be a lab exam. Learning targets here are English and Spanish.

What actually counts as one word?
Before you compare your deck to a chart, name the ruler. Most vocabulary drama collapses into three units.
| Unit | What it counts | Effect on totals | Everyday example |
|---|---|---|---|
| Word form | Each distinct spelling or inflection you see | Totals inflate fast | go, goes, went, gone = 4 |
| Lemma | Dictionary headword (citation form) | Closer to many CEFR-style orientation bands | go covers goes / went / gone in list thinking |
| Word family | Related set of derivatives | Can look larger than pure lemma counts in reading research | nation, national, nationally, nationalize |
Researchers and apps mix these units freely. That is why "B2 needs 3,000" and "B2 needs 8,000" can both show up in serious-looking posts.
Practical rule: if you cannot answer "what unit?", treat the comparison as entertainment, not measurement.
Here is the same idea as a quick visual: one spelling trail, one headword, one wider family.

Worked examples: English and Spanish
English: go and friends
One week of past-tense practice can look like dozens of "new words" if every form is a separate card.
| Counting style | Items you might count | Rough total |
|---|---|---|
| Word forms | go, goes, going, went, gone | 5 |
| Lemma | go (noun sense may split by list rules) | ~1 |
| Loose family thinking | go, going (as process), goer, outgoing… | depends how aggressive the family is |
On a lemma map, those conjugations may still sit under one headword until you add true new items (leave, arrive, depart). On a form-heavy deck, a tense drill looks like a growth spurt.
English also hides multiword pain. Look after, look up, and look into can feel like one mental verb and three separate dictionary problems. Lemma counters undercount that friction; form counters overcount spelling noise. For more on that gap, see words vs expressions in CEFR thinking.
Spanish: correr and a real paradigm
Spanish morphology makes the form/lemma gap obvious. A few tenses explode into a long form list while the dictionary still shows one verb.
| Counting style | What appears | Feel for the learner |
|---|---|---|
| Word forms | corro, corres, corre, corremos, corréis, corren, corrí, corrió, corría, corrido, corriendo… | Huge list after a few tenses |
| Lemma | correr as the catalogue entry | One item until you truly add carrera, corredor, etc. |
| Related but separate lemmas | carrera, corredor, corrido (adjective or noun senses) | Real new meanings, not just conjugation |
Agreement adds more form noise. Nuevo / nueva / nuevos / nuevas can look like four "words" in a form-based app; lemma lists often collapse them to one adjective head.
Here's the useful part: knowing correr and freezing on corrió or hayan corrido is often morphology and parsing, not a missing root. Marking every conjugation as mastered can inflate progress without adding new roots like tropezar. Practice the hard form; do not invent a second ego trophy for every person of the present.
Why "my 2000 words" is not comparable
Imagine three people. All three say "I have two thousand words." None of them share a unit.
- Person A keeps an Anki deck of 2,000 surface cards (conjugations, plurals, spelling variants). The total looks big. Coverage of real text may still lag a leaner lemma list.
- Person B tracks a frequency list of 2,000 lemmas. That is closer to CEFR-style orientation bands many maps use.
- Person C quotes a reading-research estimate framed in word families. Families bundle derivatives, so the scale is different again.
Same number. Different rulers. The first person is not automatically "ahead" of the second, and the third is not cheating - they are counting something else.
When someone says "I already have more words than B2 needs," ask: forms, lemmas, or families? Then compare to orientation on the language map, which is framed for frequent lemmas, not every spelling.

Why app catalogues rarely match research papers
Different tools do different jobs. They are not wrong. They are not the same ruler.
| Source | Job | Unit habit |
|---|---|---|
| Classroom frequency tests (X-Lex-style) | Estimate size against high-frequency lists | Often lemma / checklist style |
| Nation-style reading studies | Link coverage % to reading comfort | Often families |
| SRS apps | Schedule review of cards the user created | Whatever the user typed |
| Graded product catalogues | Progressive curriculum lists by level | Curated lemmas / entries, not pure corpus science |
Even "lemma" is not one law. List rules hide in the footnotes:
- Is email / e-mail one item or two?
- Are proper names included?
- Are multiword expressions separate entries?
- Does Spanish list se constructions separately?
- Does English split phrasal verbs?
Two honest lemma lists can still disagree by hundreds of items.
One more humility check: receptive vs productive. Recognition ("I have seen this") is cheaper than production under pressure. A catalogue mark usually means you decided this item is learned for your workflow, not that a lab verified productive control. Keep that when the percentage moves.
How this affects CEFR word charts
CEFR itself does not publish an official word quota. Public charts and apps invent or borrow orientation bands so learners have a map.
On MovaReader, the free hub uses research-informed frequent-lemma orientation (same bands as the pillar article):
| Level | Orientation range (frequent lemmas) |
|---|---|
| A1 | up to ~1,500 |
| A2 | ~1,500-2,500 |
| B1 | ~2,500-3,250 |
| B2 | ~3,250-3,750 |
| C1 | ~3,750-4,500 |
| C2 | ~4,500-5,000+ |
Read those numbers with unit discipline:
- They are navigation, not certificates.
- They sit closer to lemma thinking than to raw form dumps.
- Reading-research figures in families can look larger for the same real-world ability.
- App catalogue sizes for a level can be smaller or larger than the research band because the list is curated for product learning, not for publishing a paper.
- Spanish Plan Curricular describes teaching content; it is not one mandatory word count per level.
When a chart says "B2 = 4,000 words" and another says "B2 = 3,500 lemmas," they may be describing overlapping reality with different rulers. Details for the B2 band: How many words for B2?. Full ladder: How many words for CEFR A1-C2?.
How catalogue progress works on a language map
After the unit story, the practical question is: how do you track your English or Spanish without fake precision?
Keep the model simple and honest:
- Public map ranges show research orientation by level. Free to browse on the language map.
- Personal progress appears after you sign in, choose English or Spanish as a learning language, and mark words as learned in the catalogue.
- Progress is roughly marked items vs catalogue size for that path - learning analytics, not exact scientific coverage of all English or Spanish.
- Marks usually come from real use: reading with in-context translation, saving to a personal dictionary, then deciding an item is solid enough to mark.

What progress is not:
- not DELE / IELTS / Cambridge
- not "your official CEFR level in 60 seconds free"
- not a promise that 100% of a B2 catalogue equals fluent speaking
- not proof that two apps' "2,000 words" are equal
- not a free full CEFR exam or certificate
- not a scientific vocabulary-size lab result
MovaReader catalogues are curated progressive lists. They help you navigate A1-C2 bands. They are not identical to Milton/Alexiou research ranges, not identical to Instituto Cervantes content inventories, and not a substitute for exams.
For volume: freemium can try the loop (20 word translations/day). Practical volume of unlimited translations and dictionary work that feeds those marks sits on Reader from €1/mo. Optional spaced review after saving words lives on trainers (full unlimited trainers lean Pro). There is no free full CEFR level test. See pricing, how to use, and track vocabulary progress. Level focus after sign-in: B1 vocabulary or B2 vocabulary.
Practical advice: one consistent system beats absolute truth
You do not need a perfect universal word definition. You need a stable unit for your journey.
- Pick a primary system - catalogue marks, one SRS deck, or one frequency list - not three incompatible trophies.
- Prefer lemmas for planning CEFR bands - forms are great for grammar practice; noisy for size ego.
- Judge real pages by unknown density - if every sentence needs a lookup, comfort is low even if the total looks "B2-ish."
- Grow through reading, not only list completion - context sticks better than isolated trophy counting (build B1 vocabulary by reading).
- Re-check hard forms without double-counting ego - knowing correr does not cancel practice of corrió; you may not need a second "word" trophy for every person of the present.
- Separate English and Spanish progress - morphology load and multiword patterns differ; do not copy-paste a number between languages.
If you want one action today: open the map, treat ranges as orientation, read something slightly above comfort, and mark only items you truly own - not every glance. Coverage on real pages still wins over a prettier total (vocabulary for reading comprehension).
FAQ
Is a lemma the same as a "word" in everyday speech?
No, not exactly. Everyday "word" is fuzzy. Linguists use lemma for the headword you look up in a dictionary. Apps often mean "card" or "form." Always check which unit a number uses before you compare.
Should I count went separately from go?
For size planning against CEFR-style lemma bands, usually no. For practice, yes: irregular past forms need retrieval work. Size and practice are different jobs - do not mix the trophies.
Why does my app show more words than the map band for my level?
It may count forms, multiword cards, proper names, or a different list. Catalogue size is also curated and not identical to research orientation ranges. Compare units first, then decide if the gap is real.
Do Spanish and English need the same count method?
The three-unit problem is the same (form / lemma / family). Spanish will push more surface forms via conjugation and agreement; English will hide difficulty in phrasal verbs and spelling classes. Same map idea, different friction - keep separate progress per language.
Does marking a word "learned" prove my CEFR level?
No. It is learning analytics for your catalogue path after you sign in and choose English or Spanish. Exams and real conversation still judge performance. The map is not a free full level test or certificate.
Is 5,000 words C2?
Not automatically. ~4,500-5,000+ frequent lemmas is late-range orientation on some scales - navigation, not a diploma. Skills, text coverage, and exams still matter. See how many words for CEFR and does 5000 words mean C2?.
Next steps
- Read ranges on the free language map with lemma discipline.
- Use the pillar A1-C2 word ranges and the B2 deep-dive when you need level-specific numbers.
- Sign in, choose English or Spanish, and track one catalogue consistently while you read; mark only items you truly own.
- Optional level focus after auth: B1 vocabulary or B2 vocabulary.
- When freemium translation limits block the volume you need, Reader from €1/mo is the honest step for practical reading and dictionary use - not a certificate purchase. Method notes: how to use.
Vocabulary size is a useful mental model. It becomes noise the moment you compare incompatible units. Count like a mapmaker: same ruler, clear labels, no fake diploma.
Learn languages by reading!
Try MovaReader for just €1 — read texts with instant translation and interactive vocabulary training.
Free plan: 20 word translations/day and unlimited listening. Pro (€5/mo) unlocks full trainers.