Vocabulary Profiler
Paste any English text and every word is matched against this build's compiled word records, colored by frequency band from most common to off-list, and summarized: per-band coverage, type-token ratio, approximate lexical density, and average syllables. Band word lists can be copied or downloaded, and the report prints cleanly.
Marked-up text
- Band 1 — most common in this build
- Band 2
- Band 3
- Band 4–5 — rarest retained
- Off-list — no record in this build
Hover over a marked word for its band and, where stored, a short gloss. The band word lists below carry the same glosses for keyboard and screen-reader use.
Summary
| Band | Distinct words | Word tokens | Share of tokens |
|---|
Band word lists
About the Vocabulary Profiler
The profiler runs one server-side pass over your text. Each word token is matched against this build's compiled word records and assigned a frequency band from the record's commonness tier — a build-specific ranking band derived from the pinned source data, not a universal usage claim. Words with no record are reported as off-list rather than judged incorrect: names, inflections absent from the source lists, and non-English words all land there.
The summary reports the share of word tokens in each band, the type-token ratio (distinct words divided by total words), an approximate lexical density (the share of words outside a fixed built-in list of grammatical function words — no part-of-speech tagging of your text is performed), and the average syllable count over the words whose records carry pronunciation data.
Data sources and bounds
- Word records come from this build's compiled lexicon of pinned source lists. Membership means presence in those source lists, not validity in any particular game or register.
- Frequency bands group the build's commonness tiers: band 1 is tier 1, bands 2 and 3 match tiers 2 and 3, and band 4–5 merges the two rarest retained tiers. Tier boundaries are set per build and can change between builds.
- Tokens are runs of letters a–z with optional internal apostrophes; apostrophes are dropped for lookup. Hyphenated compounds are profiled as their parts, and accented or non-Latin characters split tokens.
- Short glosses come from this project's generated definitions corpus and are informational summaries, not authoritative dictionary entries. Many words, including most off-list tokens, have no stored gloss.
- Syllable counts come from each record's primary retained pronunciation variant; words without pronunciation data are excluded from the syllable average, and the report states how much of the text that average covers.
- Words on this build's content-policy exclusion list keep their counts but their glosses are withheld and they are left out of the band word lists and exports.
- Analysis is bounded: at most 20,000 characters and 5,000 word tokens per request. Longer text is truncated and the report says so.
- Your text is analyzed in memory to build the response and is not stored.
Results reflect data version v1/2026.07-lexical.6; records, tiers, glosses, and pronunciation coverage can change between builds.