Vocabulary Profiler

Paste any English text and every word is matched against this build's compiled word records, colored by frequency band from most common to off-list, and summarized: per-band coverage, type-token ratio, approximate lexical density, and average syllables. Band word lists can be copied or downloaded, and the report prints cleanly.

Up to 20,000 characters. Words use letters A–Z with optional internal apostrophes; everything else is left unmarked.

About the Vocabulary Profiler

The profiler runs one server-side pass over your text. Each word token is matched against this build's compiled word records and assigned a frequency band from the record's commonness tier — a build-specific ranking band derived from the pinned source data, not a universal usage claim. Words with no record are reported as off-list rather than judged incorrect: names, inflections absent from the source lists, and non-English words all land there.

The summary reports the share of word tokens in each band, the type-token ratio (distinct words divided by total words), an approximate lexical density (the share of words outside a fixed built-in list of grammatical function words — no part-of-speech tagging of your text is performed), and the average syllable count over the words whose records carry pronunciation data.

Data sources and bounds

  • Word records come from this build's compiled lexicon of pinned source lists. Membership means presence in those source lists, not validity in any particular game or register.
  • Frequency bands group the build's commonness tiers: band 1 is tier 1, bands 2 and 3 match tiers 2 and 3, and band 4–5 merges the two rarest retained tiers. Tier boundaries are set per build and can change between builds.
  • Tokens are runs of letters a–z with optional internal apostrophes; apostrophes are dropped for lookup. Hyphenated compounds are profiled as their parts, and accented or non-Latin characters split tokens.
  • Short glosses come from this project's generated definitions corpus and are informational summaries, not authoritative dictionary entries. Many words, including most off-list tokens, have no stored gloss.
  • Syllable counts come from each record's primary retained pronunciation variant; words without pronunciation data are excluded from the syllable average, and the report states how much of the text that average covers.
  • Words on this build's content-policy exclusion list keep their counts but their glosses are withheld and they are left out of the band word lists and exports.
  • Analysis is bounded: at most 20,000 characters and 5,000 word tokens per request. Longer text is truncated and the report says so.
  • Your text is analyzed in memory to build the response and is not stored.

Results reflect data version v1/2026.07-lexical.6; records, tiers, glosses, and pronunciation coverage can change between builds.