Data Sources
Loaded data version: v1/2026.07-lexical.6
· Expected release data version: 2026.07-lexical.6
· Build: 647a998abcbf
The Word Index combines several named lexical resources into a bounded search index. Coverage, ranking tiers, and pronunciations reflect those sources and their documented limitations; they are not universal facts about English. Statistics below come from the loaded compiled assets.
Coverage Statistics
Calculation contracts:
release-statistics-v1,
release-statistics-v1,
coverage-ratio-v1, and
release-statistics-v1.
These are release statistics for this exact build, not estimates of the
size or sound inventory of English.
Data Sources
We are grateful to the creators of these resources:
10000 Common Words source-membership, source-order
- Source ID / class:
google-10000-english— ranked-web-corpus-derived-list - Version:
git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8 - What it establishes: source-membership, source-order only
- Source: https://raw.githubusercontent.com/first20hours/google-100...
- License: No explicit licence grant in pinned repository metadata
- Attribution: first20hours/google-10000-english, derived from the Google Web Trillion Word Corpus and Peter Norvig's compilation
- Artifact content version:
google-10000-english.txt, SHA-256ee2d83651fbb91642bbed2bd30ead404c2cfbdfece01dacf284af6ea47795811 - Present in loaded build: yes
- Legal review: pending — The exact pinned LICENSE.md describes provenance but contains no explicit public-domain dedication or licence grant. Human review is required.
- Redistribution review: pending — The pinned repository documents provenance but does not supply an affirmative licence grant; do not describe it as public domain.
ENABLE Wordlist source-membership, surface-spelling-as-listed
- Source ID / class:
enable— historical-game-word-list - Version:
artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7 - What it establishes: source-membership, surface-spelling-as-listed only
- Source: https://www.wordgamedictionary.com/enable/download/enable.txt
- License: Licence terms not stated by the source
- Attribution: ENABLE (Enhanced North American Benchmark Lexicon); exact reuse terms are pending legal review
- Artifact content version:
enable.txt, SHA-256d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7 - Present in loaded build: yes
- Legal review: pending — The downloaded artifact is content-pinned by SHA-256, but it contains no licence notice and the configured source page does not state reuse terms.
- Redistribution review: pending — The artifact is reproducibly pinned but no affirmative reuse grant has been located; redistribution approval remains pending.
English Words (dwyl) source-membership, surface-spelling-as-listed
- Source ID / class:
dwyl-english-words— broad-orthographic-list - Version:
git:8179fe68775df3f553ef19520db065228e65d1d3 - What it establishes: source-membership, surface-spelling-as-listed only
- Source: https://raw.githubusercontent.com/dwyl/english-words/8179...
- License: Unlicense notice; upstream dataset rights unresolved
- Attribution: dwyl/english-words; the repository carries an Unlicense notice, while its pinned README attributes upstream copyright to Infochimps
- Artifact content version:
words_alpha.txt, SHA-2563ed0c94610d8bcf7c11bbb49c56aa49c7234d32b66824df91f554169e572da48 - Present in loaded build: yes
- Legal review: pending — The pinned repository licence is the Unlicense, but the pinned README says copyright in the extracted source list remains with Infochimps. Human legal review of the upstream rights chain is required.
- Redistribution review: pending — Do not assert a redistribution right until the Infochimps-to-dwyl rights chain has been reviewed by a human.
CMUdict source-membership, arpabet-pronunciation-variant
- Source ID / class:
cmudict— pronunciation-dictionary - Version:
git:74790861f652b15e4ac49015a90074ad62a27690 - What it establishes: source-membership, arpabet-pronunciation-variant only
- Source: https://raw.githubusercontent.com/cmusphinx/cmudict/74790...
- License: BSD-2-Clause
- Attribution: CMU Pronouncing Dictionary, Copyright 1993-2015 Carnegie Mellon University, BSD-2-Clause
- Artifact content version:
cmudict.dict, SHA-25681917843c7f44ce2b094ac63873c2c7a4cf802040792c455ba3ca406891c3d22 - Present in loaded build: yes
- Legal review: pending — The exact pinned commit includes BSD-2-Clause redistribution terms. Human approval of the production attribution and binary-distribution obligations is still required.
- Redistribution review: pending — BSD-2-Clause terms are present at the pin; retain copyright and licence notices. Production approval remains a recorded human decision.
Word Frequency List source-token-observation, source-token-count
- Source ID / class:
frequencywords-en-2018— subtitle-derived-token-frequency - Version:
git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018 - What it establishes: source-token-observation, source-token-count only
- Source: https://raw.githubusercontent.com/hermitdave/FrequencyWor...
- License: CC BY-SA 4.0 (content)
- Attribution: FrequencyWords content derived from OpenSubtitles, licensed CC BY-SA 4.0
- Artifact content version:
en_50k.txt, SHA-2565351ff405b1126ef555791dd4d9798a48e3e9a501a9fc481a9da957752cfb458 - Present in loaded build: yes
- Legal review: pending — The exact pinned README states MIT for code and CC BY-SA 4.0 for content. Human review is required for attribution and ShareAlike obligations on compiled and derived data.
- Redistribution review: pending — Content is declared CC BY-SA 4.0. Attribution and ShareAlike treatment of the compiled frequency evidence and derived tiers require human approval.
Enrichment provenance
Manifest SHA-256: 8258eddb1c70d47b45cf4c54a8b75fd761e2f6b36152d1b1626cab9fe6401366.
Combined source-and-content digest: 477542693badcafd96a49a013917cb986d413ee48d738561b13d04fa8f330416.
Inspect 13 enrichment artifacts Files, content seals, licences, and review state
| Dataset | Files | Content version (tree SHA-256) | Licence / legal state |
|---|---|---|---|
| word-definitions | 948 | e2ea1b8257a212308621d3612c7b3d3e55afeef624802692bb305895ef2534a9 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| morphology-tier-1 | 79 | ecfcb6f76417125af798f80f17441463e16affe80ea84f1efa3feb28000a8bc8 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| morphology-tier-2 | 316 | 5c00b2dac0ab70b3c485440e5204bb5d578558910b0ae19d7d7d56fedcf1fc70 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| cefr-tier-1 | 79 | 3ff12dd0ed357410097e07b6d931110583fd44736b3f31c86164f7681c37810b |
Reuse terms not yet approved — legal: pending; editorial: pending |
| cefr-tier-2 | 316 | 211ffddff4d6f691a7afb8185d5e42eab1fe8a6f63277d76e22a70b92602dc28 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| annotated-relations | 1 | 5f50a8b0f2869fa5cf5834ba780a9e1bc523a0e911d3a484f8401479f4331897 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| etymologies | 1 | 8770f3c83f3c1c81667bc621884eba5e913912859371831f951eea0a33abe509 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| enhanced-definitions-draft | 1 | 89c8703ce658c86d8cdfe9604e769172c84861aa15e66a21e348429311606b02 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| usage-notes-draft | 2 | bba1781e8458686d428d4dc76379f63756e57d7aa6558c1b451f54c6383d01b8 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| curated-usage-notes-draft | 1 | a7a599d9b953026f9d03f82abb527c3f4b2c468684f5c059d1dd667d277dbf5f |
Reuse terms not yet approved — legal: pending; editorial: pending |
| collocations-draft | 1 | 72477a25abae3ace65b99a99de0ce05686d3cbcefe3e268e14958c92b92964d8 |
Reuse terms not yet approved — legal: pending; editorial: pending |
| definition-work-queues | 3 | e06bc0c3d3806acfc093bf461808ab935ecb7ed555a55f5fc86f6af8116566ed |
Reuse terms not yet approved — legal: pending; editorial: pending |
| offensive-term-policy | 2 | 40ea50d53462dc19d1c7dc1327ca0ab34987358f7555da2acf9e8534b9f5f88c |
Reuse terms not yet approved — legal: pending; editorial: pending |
Field provenance
See field-by-field provenance What establishes each displayed value—and what does not
| Displayed field | Source or transformation | Limit |
|---|---|---|
| Spelling and source membership | Pinned word-list artifacts above; normalization contract v3 | Membership is source-bounded, not universal English or game validity. |
| Result order and tier | Pinned frequency/list inputs; ranking contract v2 | A reproducible search-order heuristic, not a universal usage frequency. |
| Pronunciation, syllables, rhyme keys | CMUdict artifact above plus deterministic transforms | Primarily North American coverage; pronunciations vary. |
| Definitions, examples, relations, morphology, CEFR, etymology | Enrichment manifest entries above | Legacy AI-assisted provenance is incomplete; semantic word fields require per-word editorial approval. |
| Letter counts, tile sums, and study statistics | Named analysis version over the loaded data version | Describes this compiled source set only; each study supplies its own sidecar and limitations. |
The word endpoint and result explanations cite the loaded build ID and source classes for each spelling. The compiled ledger intentionally leaves lexeme identity, inflection, part of speech, senses, and semantic relations unresolved unless separately evidenced and reviewed.
Data Processing
We process these sources to create our unified word database:
See the four processing steps Normalization, merging, ranking, and index building
| Processing Step | Description |
|---|---|
| Normalization | Lowercase and remove diacritics; reject hyphens, apostrophes, numbers, spaces, and other non-a-z tokens |
| Deduplication | Merge identical normalized spellings while retaining their source memberships |
| Tier Assignment | Classify entries 1-5 from exact source memberships, corroboration, and observed-frequency signals |
| Index Building | Create optimized indexes for fast pattern, anagram, rhyme, and prefix/suffix search |
Tier Definitions
Compare the five ranking tiers Exact build rule and product behaviour
| Tier | Deterministic build rule | How it is used |
|---|---|---|
| 1 - Observed top band | Top 10,000 eligible ENABLE members (plus A and I) actually observed in the pinned frequency corpus | Shown first in default results |
| 2 - Eligible-observation / dual-source standard band | Eligible observed ranks 10,001–25,000, plus unobserved entries present in both Google 10,000 and ENABLE | Included in default results after tier 1 |
| 3 - Extended ENABLE band | Remaining eligible observed entries and other ENABLE members, including unreliable frequency fragments demoted from the default tiers | Available when extended results are requested |
| 4 - Corroborated non-ENABLE fallback | Unobserved non-ENABLE entries found in Google 10,000 or in more than one configured word list | Hidden from default results |
| 5 - Single-source fallback | Unobserved non-ENABLE entries supported by only one broad source, plus rejected frequency fragments outside ENABLE | Hidden from default results and not given an unreviewed detail page |
These tiers are reproducible search-order heuristics, not claims about correctness, reading level, universal frequency, or game legality.
License & Usage
Each source retains its own license and attribution requirements; the links and pinned evidence above are provided for rights review. No reuse permission is asserted while that review is pending. The application code's license does not relicense third-party datasets.
Acknowledgments
We thank all contributors to the linguistic resources that make The Word Index possible, especially:
- 10000 Common Words
- ENABLE Wordlist
- English Words (dwyl)
- Word Frequency List
- CMUdict