Vocabulary Size Test

Answer about 70 quick yes/no judgments — real entries from this site's compiled word data mixed with invented spellings — and get a rough, honestly-bounded estimate of how many of this build's ranked words you recognize, plus a personal study list of the words you missed.

How it works

  1. You will see one spelling at a time. Some are real entries in this site's compiled word data; some are invented look-alikes.
  2. Answer I know it only if you could give at least one meaning for the word. Otherwise answer I don't know it.
  3. Claims about invented spellings discount your score, so honest answers give the most accurate estimate.

Takes about three minutes. Results stay in this browser only — there are no accounts and nothing is uploaded.

About the Vocabulary Size Test

The test samples real entries from this site's compiled word data across frequency-rank bands — from the most frequently observed records down to rarely observed ones — and mixes in invented spellings generated from the letter statistics of the same data. Your yes-rate on invented spellings is used to discount your yes-rate on real ones, a standard correction for over-claiming in yes/no vocabulary tests.

Methodology

  1. Real words are sampled from seven frequency-rank bands of the loaded data build (ranks 1–250 up to 25,001 and beyond), a handful per band.
  2. Each band's raw yes-rate is corrected for guessing: corrected = max(0, (raw − fa) / (1 − fa)), where fa is your yes-rate on invented spellings.
  3. Corrected rates get a monotone non-increasing fit across bands (pool-adjacent-violators, weighted by how many words each band asked), since recognizing rarer records at a higher rate than common ones is treated as sampling noise.
  4. The estimate multiplies each band's fitted rate by the band's ranked population in this build and sums. The reported range applies the same fit to 95% Wilson interval bounds per band, so small samples honestly widen the range.

Limits — read before quoting your number

  • There is no ground truth for anyone's true vocabulary. This is a rough self-assessment over one compiled data build, not a certified measurement instrument.
  • Counts are over commonness-tier 1–3 spelling records in this build — its frequency-evidenced and corroborated entries, with inflected forms counted separately. They are not deduplicated word families and not "all English words"; tiers are build-specific ranking bands, not universal usage claims.
  • Frequency ranks come from the build's pinned subtitle-derived frequency artifact, which carries its own tokenisation and corpus bias; band boundaries are build-specific, not universal usage claims.
  • Invented spellings were generated from letter statistics of this data and filtered against every compiled spelling and the offensive-word list. A generated form could still exist in some dictionary this build does not include.
  • About 70 judgments is a small sample; retakes with different item draws will move the number. Treat the range, not the midpoint, as the result.
  • Answers, results, and history never leave this browser (localStorage only). The item list itself is downloaded with answer labels, so the test is self-scored and not cheat-proof — it measures only what you honestly report.

Data sources

Real items and glosses come from the compiled lexical build shown above (spelling lists, the pinned frequency artifact, and generated definitions). Membership of a spelling in this data means only that a pinned source contains it — it is not a claim of validity in any named word game or dictionary. See About the data for the exact sources, versions, and their limits.