How We Build and Check Our Content
The Word Index's word tools and generated result pages use a compiled lexicon that we build and verify from attributed sources. This page explains where each kind of content comes from — including where we use AI assistance and how that work is reviewed.
The word database
The build combines exact, checksum-pinned artifacts from ENABLE, the dwyl English words repository, CMUdict, Google 10,000 English, and the FrequencyWords repository. Their reuse terms are not uniformly resolved: the data sources page records each version, checksum, declared licence evidence, and human legal-review state. Pending rights review keeps data-dependent pages out of search.
The primary unit is a normalized spelling record, not a declaration that we have resolved a lexeme. Surface spelling, normalized spelling, source membership, corpus observation, pronunciation variants, meanings, inflected forms, and game-list membership are separate facts. Where our sources do not distinguish homographs, senses, dialects, or inflections, the data says unresolved rather than guessing.
The five result tiers are a build-specific ranking heuristic. Observed frequency is used when the pinned frequency artifact contains a usable token; otherwise source membership and agreement provide conservative fallbacks. A tier is not proof of universal English usage, Wordle or Scrabble validity, or crossword suitability.
Why sources disagree
Each source was built for a different purpose. ENABLE is a historical game-oriented list; the broad dwyl artifact has different coverage; FrequencyWords counts subtitle-derived tokens; and CMUdict records mainly North American pronunciations. A spelling can therefore appear in one source and not another, or have several pronunciations. The product keeps those memberships and variants independent instead of reducing them to a single “valid word” flag.
Versions and reproducibility
Every release has a content-derived build ID. It binds exact source checksums, the enrichment inventory, transformation declarations, and implementation hashes. Builds retain reason-coded rejected rows, normalization collisions, duplicate conflicts, a quality report, and a release-to-release semantic diff. Changing a source or rule therefore creates a different build even if somebody forgets to rename the release.
Computed content
Word lists, letter statistics, raw tile-value sums over declared ENABLE-listed subsets, anagram groups, rhyme sets, syllable counts, are computed from the same versioned lexical observation model used by the search engine. Study analyses run in a separate reproducible build pipeline; the web application checksum-validates and renders the sealed output without recalculating its numbers. Each data study describes its own method, version, inputs, limitations, and licence-review state. Every study exposes its canonical metadata bundle, while aggregate CSV output is available only when that study's declared download gate permits it.
Definitions and AI assistance
Definitions, examples, and word-relationship notes in the source data include drafts assisted by large language models. None is approved merely because it exists. Word pages and semantic fields are published only through an explicit review record naming the reviewer, date, sources, data contract, and approved fields. The recovery registry currently starts empty, so generated dictionary content is withheld and word utility pages are noindexed until real reviews are added.
We do not claim to replace a professional dictionary. For disputed senses, etymologies, or usage advice, we recommend consulting a major dictionary; our aim is fast, practical help built on honest data.
Editorial content
Editorial explainers are controlled by a finite content manifest that records authorship, sources, dates, body checksums, corrections, and separate factual, editorial, publication, and search decisions. Study output has its own sealed registry and review gates. Machine validation may make a useful result public, but it does not imply human editorial or source-rights approval, and publication never makes either kind of content search eligible. AI-assisted drafting is recorded per item; no page claims professional lexicographic review unless a named review record supports it.
Corrections
Found something wrong — a bad definition, a missing word, a mistaken statistic? Please tell us. Corrections are recorded in the affected content or data correction history. An editorial correction can ship in an application release; a lexical fact correction follows the versioned data-build process. We would rather fix an error than defend one.