Press & Data Desk

Every published study below ships as a downloadable CSV bundle with a machine-readable metadata file, a methodology summary, explicit limitations, and exact data-version bindings in its HTTP headers. Reuse terms are recorded per study; while source-rights review is open, no blanket reuse licence is granted.

Reuse & licensing

The compiled lexicon combines named third-party sources whose redistribution rights are documented (and in several cases still under human review) in Data Sources. Because of that, reuse terms are source-bounded and per study: each row below shows the exact licence state recorded in its sealed study bundle, and nothing on this page grants rights beyond that recorded state. Downloads carry their data version, build ID, and result checksums as HTTP headers, so any figure you quote can be tied to the exact release that produced it.

When quoting a result, please attribute it to The Word Index, link the study page, and keep the study's stated limitations with the number. Results are membership and ranking facts about pinned source lists, not universal facts about English.

Published study bundles

How three English spelling lists overlap

61.97% of the safe normalized-spelling records occur in exactly one of the three pinned spelling artifacts.

Question: How much do the pinned dwyl, ENABLE, and Google 10,000 spelling lists overlap in this lexical release?

Method: Deterministic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.0).

Limitations (4)
  • The three lists have different purposes, collection methods, and scopes.
  • List membership does not establish common usage, dictionary status, game validity, or suitability for a task.
  • Normalized ASCII comparison does not evaluate senses, inflections, or regional usage.
  • Source redistribution and output reuse rights remain under human review.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'How three English spelling lists overlap' (analysis 1.0.0), https://thewordindex.com/studies/source-list-overlap

What supports The Word Index commonness tiers

Only 8.22% of safe records receive a tier directly from an eligible observed rank; the remainder use declared fallback or demotion rules.

Question: How much of the compiled lexical release has direct observed frequency evidence, and how much is ranked by bounded fallback rules?

Method: Deterministic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.0).

Limitations (4)
  • FrequencyWords is a bounded 2018 corpus-derived list and is not a timeless census of English.
  • Ranks are computed only among eligible normalized spellings after deterministic filtering.
  • Fallback source membership can prioritize results but is not direct usage evidence.
  • The study does not claim that a tier proves familiarity for an individual reader or community.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'What supports The Word Index commonness tiers' (analysis 1.0.0), https://thewordindex.com/studies/commonness-evidence

Pronunciation coverage in the compiled lexicon

The pinned pronunciation artifact covers 14.96% of the safe compiled spelling inventory; 9.12% of covered records retain more than one pronunciation variant.

Question: How completely does the pinned CMUdict artifact cover the compiled spelling population, and how often do covered spellings have variant analyses?

Method: Deterministic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); cmudict (git:74790861f652b15e4ac49015a90074ad62a27690; pronunciation-dictionary); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (4)
  • CMUdict primarily represents North American English and is not complete for every dialect, name, inflection, or specialist term.
  • Multiple entries can represent pronunciation alternatives without identifying their regional prevalence.
  • Syllable counts and rhyme keys are deterministic derivations from retained phonemes, not independent human judgments.
  • No coverage result establishes that an uncovered spelling lacks a pronunciation.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'Pronunciation coverage in the compiled lexicon' (analysis 1.0.1), https://thewordindex.com/studies/pronunciation-coverage

Letter positions by spelling length and frequency tier

For 7-letter spellings, positional leaders are S-A-R-A-I-E-S across all tiers and S-A-A-T-I-E-S in tiers 1-2.

Question: How do letter distributions change by spelling position, length, and frequency-tier cohort in this release?

Method: Deterministic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (4)
  • Each distinct normalized spelling has equal weight; the analysis is not token frequency in natural language.
  • Inflected and morphologically related spellings can each contribute a record.
  • The tier 1-2 comparison inherits the commonness model's corpus and fallback limitations.
  • Results outside the declared length window are not inferred.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'Letter positions by spelling length and frequency tier' (analysis 1.0.1), https://thewordindex.com/studies/letter-frequency

Anagram-family density in broad and higher-ranked vocabularies

Among 7-letter records, 35.62% of the full safe five-tier cohort belong to a non-singleton exact-letter-signature family, compared with 9.62% of tiers 1-2.

Question: How often does a retained spelling share its exact letter multiset with another spelling, and how does that differ between all tiers and tiers 1-2?

Method: Deterministic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (4)
  • A shared signature establishes only exact character rearrangement in this normalized build.
  • Family membership does not prove common usage, dictionary status, or game-list validity.
  • The compiled broad cohort inherits the specialist and obscure entries of its source lists.
  • The aggregate download deliberately does not redistribute the underlying spelling lists.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'Anagram-family density in broad and higher-ranked vocabularies' (analysis 1.0.1), https://thewordindex.com/studies/anagram-density

The possibility space of seven-letter honeycomb puzzles

57,581 seven-letter honeycomb-style configurations pass the v1 quality bar, drawn from 15,593 candidate alphabets over 98,009 eligible answer records in this build.

Question: How many viable seven-unique-letter honeycomb-style puzzle configurations does this lexical release support, and how are answers and pangrams distributed across them?

Method: Deterministic, build-specific analysis over enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (4)
  • Eligibility uses membership in one pinned game-history artifact and build-specific ranking tiers; it does not establish validity in any commercial word game.
  • Answer counts describe this exact normalized build and change with any source or policy update.
  • The tier 1-2 answer share is a build-specific ranking proxy, not a measured human difficulty.
  • The aggregate download lists lettersets and counts only; it does not redistribute answer word lists.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'The possibility space of seven-letter honeycomb puzzles' (analysis 1.0.1), https://thewordindex.com/studies/honeycomb-possibility-space

Word-ladder connectivity records and classic doublets

Across lengths 3-7, 51,741 eligible records form 101,904 one-letter substitution links; the longest certified shortest ladder found spans 52 steps, and 10 of 10 classic doublet pairs have a ladder in this build.

Question: How connected are this build's one-letter-substitution ladder graphs at each length, how long are the longest certified shortest ladders, and which classic doublet pairs remain solvable here?

Method: Heuristic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (4)
  • Vertex eligibility uses build-specific ranking bands and the product safety policy, so connectivity describes this exact build only.
  • The longest-ladder records are certified lower bounds from deterministic sweeps, not exhaustive all-pairs searches.
  • Classic doublet endpoints are quoted from a public-domain 1879 puzzle; historical solutions may use spellings this build excludes.
  • The aggregate download contains counts, step records, and the quoted doublet endpoints only; it does not redistribute word lists.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'Word-ladder connectivity records and classic doublets' (analysis 1.0.1), https://thewordindex.com/studies/word-ladder-records

Declared spelling-form score: a reproducible three-component formula

31.55% of tier 1-2 covered records score at least 6 points under the declared formula; the highest tier 1-2 score is 22.

Question: Which pronunciation-covered spellings receive the highest values under the declared three-component formula, and how do its score bands distribute across ranking cohorts?

Method: Heuristic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); cmudict (git:74790861f652b15e4ac49015a90074ad62a27690; pronunciation-dictionary); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (6)
  • Every weight is a fixed heuristic choice; the score is not a measured misspelling or reading-error rate.
  • Only the first retained pronunciation variant is compared; other variants and dialects are ignored.
  • CMUdict transcribes a primarily North American convention, so letter-phoneme surplus inherits its choices.
  • The declared vowel-sequence list is finite; spellings outside it can still receive points from letter-phoneme count surplus or doubled letters.
  • The vowel-sequence and doubled-letter components inspect spelling shape only; they do not compare those sequences with pronunciation.
  • Tier cohorts are build-specific ranking bands, not universal usage claims.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'Declared spelling-form score: a reproducible three-component formula' (analysis 1.0.1), https://thewordindex.com/studies/spelling-difficulty

The Delve Index: a self-audit of our machine-written definition corpus

First-party glosses document 59.98% of safe records; the 22 declared style markers occur 846 times in 2,973,912 gloss tokens.

Question: Which tier 3-5 compiled tokens occur most often in the first-party definition corpus, and how often do declared machine-prose style markers appear?

Method: Heuristic, build-specific analysis over dwyl-english-words (git:8179fe68775df3f553ef19520db065228e65d1d3; broad-orthographic-list); enable (artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7; historical-game-word-list); google-10000-english (git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8; ranked-web-corpus-derived-list); frequencywords-en-2018 (git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018; subtitle-derived-token-frequency); offensive-term-policy (2026.07-review-pending.2; first-party-content-safety-policy) (analysis version 1.0.1).

Limitations (5)
  • The glosses are first-party machine-generated text; every rate describes that corpus's writing style, not English usage.
  • Token counting ignores senses, multi-word phrases, and grammar; a marker can be the only correct word in its context.
  • Splitting on non-letters fragments possessives and abbreviations, so single-letter tokens such as 's' occur.
  • The style-marker list is a fixed editorial heuristic recorded in version control.
  • Tier joins are build-specific ranking bands, not universal usage claims.

Reuse: No reuse licence asserted; source-rights review pending (details).

Download CSV · Metadata (JSON)

Cite as: The Word Index, 'The Delve Index: a self-audit of our machine-written definition corpus' (analysis 1.0.1), https://thewordindex.com/studies/delve-index

In review (not yet downloadable)

  • Rhyme density (methodological hold) — draft (held for review).
  • Longest one-syllable spellings (methodological hold) — draft (held for review).
  • English tile-score rankings (methodological hold) — draft (held for review).
  • Homophone collision atlas — draft (held for review).
  • No perfect-rhyme partner in this build: a pronunciation-bounded census — draft (held for review).
  • Opening guesses for letter-guessing games, by exact elimination — draft (held for review).

Press contact

For interviews, custom cuts of this data, or verification of a figure, email hello@thewordindex.com. For bespoke analysis requests, see the Research Desk.