Letter positions by spelling length and frequency tier
Published for direct use; search approval pending. Its mechanically validated result is available, but the page remains noindex and outside sitemaps while human editorial and source-rights review is incomplete.
Result
For 7-letter spellings, positional leaders are S-A-R-A-I-E-S across all tiers and S-A-A-T-I-E-S in tiers 1-2.
- Analysis population
- 330,204
- All-tier position-leading letters for 7-letter records
- S-A-R-A-I-E-S
- Tier 1-2 position-leading letters for 7-letter records
- S-A-A-T-I-E-S
For 7-letter spellings, positional leaders are S-A-R-A-I-E-S across all tiers and S-A-A-T-I-E-S in tiers 1-2.
What is being counted
Each normalized spelling contributes one character to each position. Records have equal weight, so this is a lexicon-structure analysis rather than a corpus-frequency analysis. It does not use an official Wordle list.
Share held by each position's leading letter
Every leader, count, denominator, and share is in the table.
Population and exclusions
| Population stage | Records (records) |
|---|---|
| Raw lexical ledger | 394595 |
| Excluded by product policy | 150 |
| Safe normalized-spelling records | 394445 |
| Records meeting this study's inclusion criteria | 330204 |
Position-leading sequences for 7-letter spellings
| Ranking cohort | Spelling length | Position-leading letters (letter sequence) | Records at this length (records) |
|---|---|---|---|
| all-tiers | 7 | S-A-R-A-I-E-S | 42758 |
| tiers-1-2 | 7 | S-A-A-T-I-E-S | 4264 |
The leading letter at each 7-letter position
| Position | All-tier leader (letter) | All-tier leader count (records) | All-tier records (records) | All-tier leader share (percent) | Tier 1-2 leader (letter) | Tier 1-2 leader count (records) | Tier 1-2 records (records) | Tier 1-2 leader share (percent) |
|---|---|---|---|---|---|---|---|---|
| 1 | S | 4624 | 42758 | 10.81 | S | 561 | 4264 | 13.16 |
| 2 | A | 7247 | 42758 | 16.95 | A | 681 | 4264 | 15.97 |
| 3 | R | 4207 | 42758 | 9.84 | A | 481 | 4264 | 11.28 |
| 4 | A | 3522 | 42758 | 8.24 | T | 429 | 4264 | 10.06 |
| 5 | I | 8113 | 42758 | 18.97 | I | 978 | 4264 | 22.94 |
| 6 | E | 9677 | 42758 | 22.63 | E | 1213 | 4264 | 28.45 |
| 7 | S | 8944 | 42758 | 20.92 | S | 1014 | 4264 | 23.78 |
Methodology
Research question: How do letter distributions change by spelling position, length, and frequency-tier cohort in this release?
Why it is useful: The study distinguishes patterns in the broad five-tier cohort from patterns in its higher-ranked tier 1-2 cohort.
Included
letter-frequency.length-3-to-12: Include safe normalized ASCII spellings of lengths 3 through 12, grouped into all tiers and tiers 1-2.
Excluded
letter-frequency.product-policy: Exclude spellings matched by the exact first-party content-safety policy before aggregation.letter-frequency.outside-length-window: Exclude spellings shorter than 3 or longer than 12 from positional comparisons.
Transformations
letter-frequency.position-v1: Enumerate one-indexed character positions and aggregate exact letters separately by length and tier cohort.
Calculations
research.count-v1- Count spelling and positional-letter observations in each declared cohort. Formula:
count(records). research.share-v1- Calculate a letter's exact share at a declared position and length. Formula:
100 * letter_count / cohort_spelling_count. research.positional-leader-v1- Select the highest-count letter independently at each position, breaking equal-count ties by ascending ASCII letter. Formula:
argmax(letter_count, tie_break=letter_ascending).
Limitations
- Each distinct normalized spelling has equal weight; the analysis is not token frequency in natural language.
- Inflected and morphologically related spellings can each contribute a record.
- The tier 1-2 comparison inherits the commonness model's corpus and fallback limitations.
- Results outside the declared length window are not inferred.
Versions, sources, and review
dwyl-english-words git:8179fe68775df3f553ef19520db065228e65d1d3
source-membership; surface-spelling-as-listed.
Licence: Unlicense notice; upstream dataset rights unresolved. Legal review: pending. Redistribution review: pending.
- The pinned repository licence is the Unlicense, but the pinned README says copyright in the extracted source list remains with Infochimps. Human legal review of the upstream rights chain is required.
- Do not assert a redistribution right until the Infochimps-to-dwyl rights chain has been reviewed by a human.
- Dialect or scope: unspecified
dwyl/english-words; the repository carries an Unlicense notice, while its pinned README attributes upstream copyright to Infochimps
enable artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7
source-membership; surface-spelling-as-listed.
Licence: Licence terms not stated by the source. Legal review: pending. Redistribution review: pending.
- The downloaded artifact is content-pinned by SHA-256, but it contains no licence notice and the configured source page does not state reuse terms.
- The artifact is reproducibly pinned but no affirmative reuse grant has been located; redistribution approval remains pending.
- Dialect or scope: North-American-oriented; exact edition metadata unavailable
ENABLE (Enhanced North American Benchmark Lexicon); exact reuse terms are pending legal review
google-10000-english git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8
source-membership; source-order.
Licence: No explicit licence grant in pinned repository metadata. Legal review: pending. Redistribution review: pending.
- The exact pinned LICENSE.md describes provenance but contains no explicit public-domain dedication or licence grant. Human review is required.
- The pinned repository documents provenance but does not supply an affirmative licence grant; do not describe it as public domain.
- Dialect or scope: USA no-swears variant
first20hours/google-10000-english, derived from the Google Web Trillion Word Corpus and Peter Norvig's compilation
frequencywords-en-2018 git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018
source-token-observation; source-token-count.
Licence: CC BY-SA 4.0 (content). Legal review: pending. Redistribution review: pending.
- The exact pinned README states MIT for code and CC BY-SA 4.0 for content. Human review is required for attribution and ShareAlike obligations on compiled and derived data.
- Content is declared CC BY-SA 4.0. Attribution and ShareAlike treatment of the compiled frequency evidence and derived tiers require human approval.
- Dialect or scope: unspecified
FrequencyWords content derived from OpenSubtitles, licensed CC BY-SA 4.0
offensive-term-policy 2026.07-review-pending.2
content_filter; search_results; result_counts.
Licence: Reuse terms not yet approved. Legal review: pending. Redistribution review: pending.
- This bounded policy list cannot determine whether every usage is offensive or benign.
The Word Index first-party content-safety exclusion policy
Citations
- The Word Index pinned lexical source manifest, The Word Index, version 2026.07-lexical.6 (dataset).
- FrequencyWords English 2018, FrequencyWords, version git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018 (dataset).
- The Word Index content-safety exclusion policy, The Word Index, version 2026.07-review-pending.2 (policy).
Data and reproducibility
Source-rights review is pending. The available file contains aggregate computed measures, not a redistributed spelling list. Availability does not grant broader reuse rights; consult the source and licence register before reuse.
Download the canonical aggregate CSV
View the canonical JSON study bundle.
Reproduce this version from a clean checkout:
make studies_reproduce STUDY=letter-frequency