Reproducible Word Data Studies
Original analyses generated from the same pinned lexical release that powers the product. Every displayed number, table, chart, and aggregate download is projected from one checksum-bound artifact.
Initial published cohort
These are the strongest source-bounded questions the current data can answer mechanically. Publication does not grant search eligibility: every study remains noindex and absent from sitemaps pending human editorial and source-rights review.
How three English spelling lists overlap
61.97% of the safe normalized-spelling records occur in exactly one of the three pinned spelling artifacts.
Inspect study → Artifact validatedWhat supports The Word Index commonness tiers
Only 8.22% of safe records receive a tier directly from an eligible observed rank; the remainder use declared fallback or demotion rules.
Inspect study → Artifact validatedPronunciation coverage in the compiled lexicon
The pinned pronunciation artifact covers 14.96% of the safe compiled spelling inventory; 9.12% of covered records retain more than one pronunciation variant.
Inspect study → Artifact validatedLetter positions by spelling length and frequency tier
For 7-letter spellings, positional leaders are S-A-R-A-I-E-S across all tiers and S-A-A-T-I-E-S in tiers 1-2.
Inspect study → Artifact validatedAnagram-family density in broad and higher-ranked vocabularies
Among 7-letter records, 35.62% of the full safe five-tier cohort belong to a non-singleton exact-letter-signature family, compared with 9.62% of tiers 1-2.
Inspect study → Artifact validatedThe possibility space of seven-letter honeycomb puzzles
57,581 seven-letter honeycomb-style configurations pass the v1 quality bar, drawn from 15,593 candidate alphabets over 98,009 eligible answer records in this build.
Inspect study → Artifact validatedWord-ladder connectivity records and classic doublets
Across lengths 3-7, 51,741 eligible records form 101,904 one-letter substitution links; the longest certified shortest ladder found spans 52 steps, and 10 of 10 classic doublet pairs have a ladder in this build.
Inspect study → Artifact validatedDeclared spelling-form score: a reproducible three-component formula
31.55% of tier 1-2 covered records score at least 6 points under the declared formula; the highest tier 1-2 score is 22.
Inspect study → Artifact validatedThe Delve Index: a self-audit of our machine-written definition corpus
First-party glosses document 59.98% of safe records; the 22 declared style markers occur 846 times in 2,973,912 gloss tokens.
Inspect study →Research review queue
These ideas have identifiable user relevance, but the present data or method does not yet support publication-quality conclusions. Candidate generation never publishes or indexes a study.
- Rhyme density (methodological hold) — Legacy study held with no computed result pending methodological development, factual validation, and source-rights review.
- Longest one-syllable spellings (methodological hold) — Legacy study held with no computed result pending methodological development, source-scope validation, factual validation, and source-rights review.
- English tile-score rankings (methodological hold) — Legacy game-specific claims remain withheld pending rules validation, factual validation, trademark review, and source-rights review.
- Homophone collision atlas — New computed census output remains draft and noindex pending factual validation, source-rights review, and separate publication and search decisions.
- No perfect-rhyme partner in this build: a pronunciation-bounded census — New computed census output remains draft and noindex pending factual validation, source-rights review, and separate publication and search decisions.
- Opening guesses for letter-guessing games, by exact elimination — New computed strategy tables remain draft and noindex pending factual validation, source-rights review, and separate publication and search decisions.
How publication works
A finite version-controlled registry defines the question, sources, criteria, calculations, limitations, review state, and reproducibility command. The build creates canonical JSON and CSV once. The web app verifies their checksums and exact lexical-release binding; it never recalculates a missing result.
Current study release: 9d4a107b7ae194af.
Review the lexical source and licence register and the wider data methodology.