Reproducible Word Data Studies

Original analyses generated from the same pinned lexical release that powers the product. Every displayed number, table, chart, and aggregate download is projected from one checksum-bound artifact.

Initial published cohort

These are the strongest source-bounded questions the current data can answer mechanically. Publication does not grant search eligibility: every study remains noindex and absent from sitemaps pending human editorial and source-rights review.

Artifact validated

How three English spelling lists overlap

61.97% of the safe normalized-spelling records occur in exactly one of the three pinned spelling artifacts.

Inspect study →
Artifact validated

What supports The Word Index commonness tiers

Only 8.22% of safe records receive a tier directly from an eligible observed rank; the remainder use declared fallback or demotion rules.

Inspect study →
Artifact validated

Pronunciation coverage in the compiled lexicon

The pinned pronunciation artifact covers 14.96% of the safe compiled spelling inventory; 9.12% of covered records retain more than one pronunciation variant.

Inspect study →
Artifact validated

Letter positions by spelling length and frequency tier

For 7-letter spellings, positional leaders are S-A-R-A-I-E-S across all tiers and S-A-A-T-I-E-S in tiers 1-2.

Inspect study →
Artifact validated

Anagram-family density in broad and higher-ranked vocabularies

Among 7-letter records, 35.62% of the full safe five-tier cohort belong to a non-singleton exact-letter-signature family, compared with 9.62% of tiers 1-2.

Inspect study →
Artifact validated

The possibility space of seven-letter honeycomb puzzles

57,581 seven-letter honeycomb-style configurations pass the v1 quality bar, drawn from 15,593 candidate alphabets over 98,009 eligible answer records in this build.

Inspect study →
Artifact validated

Word-ladder connectivity records and classic doublets

Across lengths 3-7, 51,741 eligible records form 101,904 one-letter substitution links; the longest certified shortest ladder found spans 52 steps, and 10 of 10 classic doublet pairs have a ladder in this build.

Inspect study →
Artifact validated

Declared spelling-form score: a reproducible three-component formula

31.55% of tier 1-2 covered records score at least 6 points under the declared formula; the highest tier 1-2 score is 22.

Inspect study →
Artifact validated

The Delve Index: a self-audit of our machine-written definition corpus

First-party glosses document 59.98% of safe records; the 22 declared style markers occur 846 times in 2,973,912 gloss tokens.

Inspect study →

Research review queue

These ideas have identifiable user relevance, but the present data or method does not yet support publication-quality conclusions. Candidate generation never publishes or indexes a study.

  • Rhyme density (methodological hold) — Legacy study held with no computed result pending methodological development, factual validation, and source-rights review.
  • Longest one-syllable spellings (methodological hold) — Legacy study held with no computed result pending methodological development, source-scope validation, factual validation, and source-rights review.
  • English tile-score rankings (methodological hold) — Legacy game-specific claims remain withheld pending rules validation, factual validation, trademark review, and source-rights review.
  • Homophone collision atlas — New computed census output remains draft and noindex pending factual validation, source-rights review, and separate publication and search decisions.
  • No perfect-rhyme partner in this build: a pronunciation-bounded census — New computed census output remains draft and noindex pending factual validation, source-rights review, and separate publication and search decisions.
  • Opening guesses for letter-guessing games, by exact elimination — New computed strategy tables remain draft and noindex pending factual validation, source-rights review, and separate publication and search decisions.

How publication works

A finite version-controlled registry defines the question, sources, criteria, calculations, limitations, review state, and reproducibility command. The build creates canonical JSON and CSV once. The web app verifies their checksums and exact lexical-release binding; it never recalculates a missing result.

Current study release: 9d4a107b7ae194af.

Review the lexical source and licence register and the wider data methodology.