Data Sources

Loaded data version: v1/2026.07-lexical.6 · Expected release data version: 2026.07-lexical.6 · Build: 647a998abcbf

Human source-rights review is pending. The exact evidence and unresolved questions are listed below. Data-dependent pages remain excluded from search, and no reuse licence for compiled study results is granted while this review is open.

The Word Index combines several named lexical resources into a bounded search index. Coverage, ranking tiers, and pronunciations reflect those sources and their documented limitations; they are not universal facts about English. Statistics below come from the loaded compiled assets.

Coverage Statistics

394,595 Normalized spelling records
59,108 Spelling records with pronunciation evidence
15% Compiled spelling coverage
18,001 Distinct indexed rhyme keys
117,486 Normalized CMUdict source headwords

Calculation contracts: release-statistics-v1, release-statistics-v1, coverage-ratio-v1, and release-statistics-v1. These are release statistics for this exact build, not estimates of the size or sound inventory of English.

Data Sources

We are grateful to the creators of these resources:

10000 Common Words source-membership, source-order
  • Source ID / class: google-10000-english — ranked-web-corpus-derived-list
  • Version: git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8
  • What it establishes: source-membership, source-order only
  • Source: https://raw.githubusercontent.com/first20hours/google-100...
  • License: No explicit licence grant in pinned repository metadata
  • Attribution: first20hours/google-10000-english, derived from the Google Web Trillion Word Corpus and Peter Norvig's compilation
  • Artifact content version: google-10000-english.txt, SHA-256 ee2d83651fbb91642bbed2bd30ead404c2cfbdfece01dacf284af6ea47795811
  • Present in loaded build: yes
  • Legal review: pending — The exact pinned LICENSE.md describes provenance but contains no explicit public-domain dedication or licence grant. Human review is required.
  • Redistribution review: pending — The pinned repository documents provenance but does not supply an affirmative licence grant; do not describe it as public domain.
ENABLE Wordlist source-membership, surface-spelling-as-listed
  • Source ID / class: enable — historical-game-word-list
  • Version: artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7
  • What it establishes: source-membership, surface-spelling-as-listed only
  • Source: https://www.wordgamedictionary.com/enable/download/enable.txt
  • License: Licence terms not stated by the source
  • Attribution: ENABLE (Enhanced North American Benchmark Lexicon); exact reuse terms are pending legal review
  • Artifact content version: enable.txt, SHA-256 d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7
  • Present in loaded build: yes
  • Legal review: pending — The downloaded artifact is content-pinned by SHA-256, but it contains no licence notice and the configured source page does not state reuse terms.
  • Redistribution review: pending — The artifact is reproducibly pinned but no affirmative reuse grant has been located; redistribution approval remains pending.
English Words (dwyl) source-membership, surface-spelling-as-listed
  • Source ID / class: dwyl-english-words — broad-orthographic-list
  • Version: git:8179fe68775df3f553ef19520db065228e65d1d3
  • What it establishes: source-membership, surface-spelling-as-listed only
  • Source: https://raw.githubusercontent.com/dwyl/english-words/8179...
  • License: Unlicense notice; upstream dataset rights unresolved
  • Attribution: dwyl/english-words; the repository carries an Unlicense notice, while its pinned README attributes upstream copyright to Infochimps
  • Artifact content version: words_alpha.txt, SHA-256 3ed0c94610d8bcf7c11bbb49c56aa49c7234d32b66824df91f554169e572da48
  • Present in loaded build: yes
  • Legal review: pending — The pinned repository licence is the Unlicense, but the pinned README says copyright in the extracted source list remains with Infochimps. Human legal review of the upstream rights chain is required.
  • Redistribution review: pending — Do not assert a redistribution right until the Infochimps-to-dwyl rights chain has been reviewed by a human.
CMUdict source-membership, arpabet-pronunciation-variant
  • Source ID / class: cmudict — pronunciation-dictionary
  • Version: git:74790861f652b15e4ac49015a90074ad62a27690
  • What it establishes: source-membership, arpabet-pronunciation-variant only
  • Source: https://raw.githubusercontent.com/cmusphinx/cmudict/74790...
  • License: BSD-2-Clause
  • Attribution: CMU Pronouncing Dictionary, Copyright 1993-2015 Carnegie Mellon University, BSD-2-Clause
  • Artifact content version: cmudict.dict, SHA-256 81917843c7f44ce2b094ac63873c2c7a4cf802040792c455ba3ca406891c3d22
  • Present in loaded build: yes
  • Legal review: pending — The exact pinned commit includes BSD-2-Clause redistribution terms. Human approval of the production attribution and binary-distribution obligations is still required.
  • Redistribution review: pending — BSD-2-Clause terms are present at the pin; retain copyright and licence notices. Production approval remains a recorded human decision.
Word Frequency List source-token-observation, source-token-count
  • Source ID / class: frequencywords-en-2018 — subtitle-derived-token-frequency
  • Version: git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018
  • What it establishes: source-token-observation, source-token-count only
  • Source: https://raw.githubusercontent.com/hermitdave/FrequencyWor...
  • License: CC BY-SA 4.0 (content)
  • Attribution: FrequencyWords content derived from OpenSubtitles, licensed CC BY-SA 4.0
  • Artifact content version: en_50k.txt, SHA-256 5351ff405b1126ef555791dd4d9798a48e3e9a501a9fc481a9da957752cfb458
  • Present in loaded build: yes
  • Legal review: pending — The exact pinned README states MIT for code and CC BY-SA 4.0 for content. Human review is required for attribution and ShareAlike obligations on compiled and derived data.
  • Redistribution review: pending — Content is declared CC BY-SA 4.0. Attribution and ShareAlike treatment of the compiled frequency evidence and derived tiers require human approval.

Enrichment provenance

Manifest SHA-256: 8258eddb1c70d47b45cf4c54a8b75fd761e2f6b36152d1b1626cab9fe6401366. Combined source-and-content digest: 477542693badcafd96a49a013917cb986d413ee48d738561b13d04fa8f330416.

Inspect 13 enrichment artifacts Files, content seals, licences, and review state
Enrichment artifacts included in the loaded lexical build.
DatasetFilesContent version (tree SHA-256)Licence / legal state
word-definitions 948 e2ea1b8257a212308621d3612c7b3d3e55afeef624802692bb305895ef2534a9 Reuse terms not yet approved — legal: pending; editorial: pending
morphology-tier-1 79 ecfcb6f76417125af798f80f17441463e16affe80ea84f1efa3feb28000a8bc8 Reuse terms not yet approved — legal: pending; editorial: pending
morphology-tier-2 316 5c00b2dac0ab70b3c485440e5204bb5d578558910b0ae19d7d7d56fedcf1fc70 Reuse terms not yet approved — legal: pending; editorial: pending
cefr-tier-1 79 3ff12dd0ed357410097e07b6d931110583fd44736b3f31c86164f7681c37810b Reuse terms not yet approved — legal: pending; editorial: pending
cefr-tier-2 316 211ffddff4d6f691a7afb8185d5e42eab1fe8a6f63277d76e22a70b92602dc28 Reuse terms not yet approved — legal: pending; editorial: pending
annotated-relations 1 5f50a8b0f2869fa5cf5834ba780a9e1bc523a0e911d3a484f8401479f4331897 Reuse terms not yet approved — legal: pending; editorial: pending
etymologies 1 8770f3c83f3c1c81667bc621884eba5e913912859371831f951eea0a33abe509 Reuse terms not yet approved — legal: pending; editorial: pending
enhanced-definitions-draft 1 89c8703ce658c86d8cdfe9604e769172c84861aa15e66a21e348429311606b02 Reuse terms not yet approved — legal: pending; editorial: pending
usage-notes-draft 2 bba1781e8458686d428d4dc76379f63756e57d7aa6558c1b451f54c6383d01b8 Reuse terms not yet approved — legal: pending; editorial: pending
curated-usage-notes-draft 1 a7a599d9b953026f9d03f82abb527c3f4b2c468684f5c059d1dd667d277dbf5f Reuse terms not yet approved — legal: pending; editorial: pending
collocations-draft 1 72477a25abae3ace65b99a99de0ce05686d3cbcefe3e268e14958c92b92964d8 Reuse terms not yet approved — legal: pending; editorial: pending
definition-work-queues 3 e06bc0c3d3806acfc093bf461808ab935ecb7ed555a55f5fc86f6af8116566ed Reuse terms not yet approved — legal: pending; editorial: pending
offensive-term-policy 2 40ea50d53462dc19d1c7dc1327ca0ab34987358f7555da2acf9e8534b9f5f88c Reuse terms not yet approved — legal: pending; editorial: pending

Field provenance

See field-by-field provenance What establishes each displayed value—and what does not
Provenance and limitations for fields displayed by the product.
Displayed fieldSource or transformationLimit
Spelling and source membershipPinned word-list artifacts above; normalization contract v3Membership is source-bounded, not universal English or game validity.
Result order and tierPinned frequency/list inputs; ranking contract v2A reproducible search-order heuristic, not a universal usage frequency.
Pronunciation, syllables, rhyme keysCMUdict artifact above plus deterministic transformsPrimarily North American coverage; pronunciations vary.
Definitions, examples, relations, morphology, CEFR, etymologyEnrichment manifest entries aboveLegacy AI-assisted provenance is incomplete; semantic word fields require per-word editorial approval.
Letter counts, tile sums, and study statisticsNamed analysis version over the loaded data versionDescribes this compiled source set only; each study supplies its own sidecar and limitations.

The word endpoint and result explanations cite the loaded build ID and source classes for each spelling. The compiled ledger intentionally leaves lexeme identity, inflection, part of speech, senses, and semantic relations unresolved unless separately evidenced and reviewed.

Data Processing

We process these sources to create our unified word database:

See the four processing steps Normalization, merging, ranking, and index building
Deterministic processing steps used to compile the lexical indexes.
Processing Step Description
Normalization Lowercase and remove diacritics; reject hyphens, apostrophes, numbers, spaces, and other non-a-z tokens
Deduplication Merge identical normalized spellings while retaining their source memberships
Tier Assignment Classify entries 1-5 from exact source memberships, corroboration, and observed-frequency signals
Index Building Create optimized indexes for fast pattern, anagram, rhyme, and prefix/suffix search

Tier Definitions

Compare the five ranking tiers Exact build rule and product behaviour
Build-specific ranking tiers and their product behavior.
Tier Deterministic build rule How it is used
1 - Observed top band Top 10,000 eligible ENABLE members (plus A and I) actually observed in the pinned frequency corpus Shown first in default results
2 - Eligible-observation / dual-source standard band Eligible observed ranks 10,001–25,000, plus unobserved entries present in both Google 10,000 and ENABLE Included in default results after tier 1
3 - Extended ENABLE band Remaining eligible observed entries and other ENABLE members, including unreliable frequency fragments demoted from the default tiers Available when extended results are requested
4 - Corroborated non-ENABLE fallback Unobserved non-ENABLE entries found in Google 10,000 or in more than one configured word list Hidden from default results
5 - Single-source fallback Unobserved non-ENABLE entries supported by only one broad source, plus rejected frequency fragments outside ENABLE Hidden from default results and not given an unreviewed detail page

These tiers are reproducible search-order heuristics, not claims about correctness, reading level, universal frequency, or game legality.

License & Usage

Each source retains its own license and attribution requirements; the links and pinned evidence above are provided for rights review. No reuse permission is asserted while that review is pending. The application code's license does not relicense third-party datasets.

Acknowledgments

We thank all contributors to the linguistic resources that make The Word Index possible, especially:

  • 10000 Common Words
  • ENABLE Wordlist
  • English Words (dwyl)
  • Word Frequency List
  • CMUdict