Anagram-family density in broad and higher-ranked vocabularies

Published for direct use; search approval pending. Its mechanically validated result is available, but the page remains noindex and outside sitemaps while human editorial and source-rights review is incomplete.

Result

Among 7-letter records, 35.62% of the full safe five-tier cohort belong to a non-singleton exact-letter-signature family, compared with 9.62% of tiers 1-2.

Analysis population
330,204
All-tier 7-letter records in non-singleton anagram families
35.62%
Tier 1-2 7-letter records in non-singleton anagram families
9.62%

Among 7-letter records, 35.62% of the full safe five-tier cohort belong to a non-singleton exact-letter-signature family, compared with 9.62% of tiers 1-2.

What an anagram family means here

Spellings are grouped only when their normalized ASCII character multisets are identical. Family membership describes this exact build; it does not establish universal English or game-list validity.

Share of records with at least one exact anagram

Exact cohort counts and percentages appear in the table.

Read the accessible data table.

Population and exclusions

The exact population used by this version of the analysis.
Population stage Records (records)
Raw lexical ledger 394595
Excluded by product policy 150
Safe normalized-spelling records 394445
Records meeting this study's inclusion criteria 330204

Anagram density by length and ranking cohort

A non-singleton family contains at least two normalized spellings with the same exact letter multiset.
Spelling length All-tier records (records) Tier 1-2 records (records) All-tier records in non-singleton families (records) Tier 1-2 records in non-singleton families (records) All-tier records in non-singleton families (percent) Tier 1-2 records in non-singleton families (percent) Largest all-tier family (records) Largest tier 1-2 family (records)
3 2290 709 1607 291 70.17 41.04 6 5
4 7320 2067 4764 793 65.08 38.36 9 5
5 16135 3160 9108 819 56.45 25.92 14 7
6 30344 4035 14311 665 47.16 16.48 15 4
7 42758 4264 15230 410 35.62 9.62 14 4
8 52746 3698 12927 168 24.51 4.54 9 5
9 54755 2887 8951 54 16.35 1.87 10 2
10 49834 1918 4796 18 9.62 0.94 5 2
11 41449 1142 2294 8 5.53 0.70 6 2
12 32573 616 1224 2 3.76 0.32 4 2

Family-size distribution

Aggregate family counts without a downloadable spelling list.
Spelling length Ranking cohort Spellings per signature Signature families (families) Spelling memberships (records)
3 all-tiers 1 683 683
3 all-tiers 2 387 774
3 all-tiers 3 182 546
3 all-tiers 4 54 216
3 all-tiers 5 13 65
3 all-tiers 6 1 6
3 tiers-1-2 1 418 418
3 tiers-1-2 2 123 246
3 tiers-1-2 3 12 36
3 tiers-1-2 4 1 4
3 tiers-1-2 5 1 5
4 all-tiers 1 2556 2556
4 all-tiers 2 913 1826
4 all-tiers 3 412 1236
4 all-tiers 4 191 764
4 all-tiers 5 94 470
4 all-tiers 6 38 228
4 all-tiers 7 20 140
4 all-tiers 8 8 64
4 all-tiers 9 4 36
4 tiers-1-2 1 1274 1274
4 tiers-1-2 2 265 530
4 tiers-1-2 3 59 177
4 tiers-1-2 4 19 76
4 tiers-1-2 5 2 10
5 all-tiers 1 7027 7027
5 all-tiers 2 1922 3844
5 all-tiers 3 717 2151
5 all-tiers 4 317 1268
5 all-tiers 5 142 710
5 all-tiers 6 80 480
5 all-tiers 7 34 238
5 all-tiers 8 15 120
5 all-tiers 9 15 135
5 all-tiers 10 4 40
5 all-tiers 11 4 44
5 all-tiers 12 2 24
5 all-tiers 13 2 26
5 all-tiers 14 2 28
5 tiers-1-2 1 2341 2341
5 tiers-1-2 2 298 596
5 tiers-1-2 3 51 153
5 tiers-1-2 4 12 48
5 tiers-1-2 5 3 15
5 tiers-1-2 7 1 7
6 all-tiers 1 16033 16033
6 all-tiers 2 3482 6964
6 all-tiers 3 1122 3366
6 all-tiers 4 453 1812
6 all-tiers 5 197 985
6 all-tiers 6 86 516
6 all-tiers 7 46 322
6 all-tiers 8 20 160
6 all-tiers 9 11 99
6 all-tiers 10 5 50
6 all-tiers 11 2 22
6 all-tiers 15 1 15
6 tiers-1-2 1 3370 3370
6 tiers-1-2 2 260 520
6 tiers-1-2 3 43 129
6 tiers-1-2 4 4 16
7 all-tiers 1 27528 27528
7 all-tiers 2 4426 8852
7 all-tiers 3 1136 3408
7 all-tiers 4 384 1536
7 all-tiers 5 134 670
7 all-tiers 6 56 336
7 all-tiers 7 32 224
7 all-tiers 8 11 88
7 all-tiers 9 8 72
7 all-tiers 10 3 30
7 all-tiers 14 1 14
7 tiers-1-2 1 3854 3854
7 tiers-1-2 2 168 336
7 tiers-1-2 3 22 66
7 tiers-1-2 4 2 8
8 all-tiers 1 39819 39819
8 all-tiers 2 4435 8870
8 all-tiers 3 844 2532
8 all-tiers 4 226 904
8 all-tiers 5 67 335
8 all-tiers 6 24 144
8 all-tiers 7 12 84
8 all-tiers 8 5 40
8 all-tiers 9 2 18
8 tiers-1-2 1 3530 3530
8 tiers-1-2 2 77 154
8 tiers-1-2 3 3 9
8 tiers-1-2 5 1 5
9 all-tiers 1 45804 45804
9 all-tiers 2 3386 6772
9 all-tiers 3 532 1596
9 all-tiers 4 95 380
9 all-tiers 5 29 145
9 all-tiers 6 8 48
9 all-tiers 10 1 10
9 tiers-1-2 1 2833 2833
9 tiers-1-2 2 27 54
10 all-tiers 1 45038 45038
10 all-tiers 2 2012 4024
10 all-tiers 3 207 621
10 all-tiers 4 34 136
10 all-tiers 5 3 15
10 tiers-1-2 1 1900 1900
10 tiers-1-2 2 9 18
11 all-tiers 1 39155 39155
11 all-tiers 2 1044 2088
11 all-tiers 3 58 174
11 all-tiers 4 4 16
11 all-tiers 5 2 10
11 all-tiers 6 1 6
11 tiers-1-2 1 1134 1134
11 tiers-1-2 2 4 8
12 all-tiers 1 31349 31349
12 all-tiers 2 571 1142
12 all-tiers 3 26 78
12 all-tiers 4 1 4
12 tiers-1-2 1 614 614
12 tiers-1-2 2 1 2

Methodology

Research question: How often does a retained spelling share its exact letter multiset with another spelling, and how does that differ between all tiers and tiers 1-2?

Why it is useful: The comparison shows how apparent anagram choice differs between the broad five-tier cohort and its tier 1-2 subset; neither cohort establishes universal or game-list validity.

Included

  • anagram-density.length-3-to-12: Include safe normalized ASCII spellings of lengths 3 through 12 and compare all tiers with tiers 1-2.

Excluded

  • anagram-density.product-policy: Exclude spellings matched by the exact first-party content-safety policy before grouping.
  • anagram-density.outside-length-window: Exclude spellings outside lengths 3 through 12 from the study population.

Transformations

  1. anagram-density.signature-v1: Sort each spelling's characters into an exact multiset signature and aggregate family sizes by length and cohort.

Calculations

research.count-v1
Count spellings, signatures, or spelling memberships in each declared family-size group. Formula: count(records).
research.share-v1
Calculate the exact share of spellings belonging to a non-singleton anagram family. Formula: 100 * records_in_families_size_ge_2 / cohort_records.
research.maximum-v1
Take the largest exact-signature family size within a declared length and ranking cohort. Formula: max(family_size).

Limitations

  • A shared signature establishes only exact character rearrangement in this normalized build.
  • Family membership does not prove common usage, dictionary status, or game-list validity.
  • The compiled broad cohort inherits the specialist and obscure entries of its source lists.
  • The aggregate download deliberately does not redistribute the underlying spelling lists.

Versions, sources, and review

dwyl-english-words git:8179fe68775df3f553ef19520db065228e65d1d3

source-membership; surface-spelling-as-listed.

Licence: Unlicense notice; upstream dataset rights unresolved. Legal review: pending. Redistribution review: pending.

  • The pinned repository licence is the Unlicense, but the pinned README says copyright in the extracted source list remains with Infochimps. Human legal review of the upstream rights chain is required.
  • Do not assert a redistribution right until the Infochimps-to-dwyl rights chain has been reviewed by a human.
  • Dialect or scope: unspecified

dwyl/english-words; the repository carries an Unlicense notice, while its pinned README attributes upstream copyright to Infochimps

enable artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7

source-membership; surface-spelling-as-listed.

Licence: Licence terms not stated by the source. Legal review: pending. Redistribution review: pending.

  • The downloaded artifact is content-pinned by SHA-256, but it contains no licence notice and the configured source page does not state reuse terms.
  • The artifact is reproducibly pinned but no affirmative reuse grant has been located; redistribution approval remains pending.
  • Dialect or scope: North-American-oriented; exact edition metadata unavailable

ENABLE (Enhanced North American Benchmark Lexicon); exact reuse terms are pending legal review

google-10000-english git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8

source-membership; source-order.

Licence: No explicit licence grant in pinned repository metadata. Legal review: pending. Redistribution review: pending.

  • The exact pinned LICENSE.md describes provenance but contains no explicit public-domain dedication or licence grant. Human review is required.
  • The pinned repository documents provenance but does not supply an affirmative licence grant; do not describe it as public domain.
  • Dialect or scope: USA no-swears variant

first20hours/google-10000-english, derived from the Google Web Trillion Word Corpus and Peter Norvig's compilation

frequencywords-en-2018 git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018

source-token-observation; source-token-count.

Licence: CC BY-SA 4.0 (content). Legal review: pending. Redistribution review: pending.

  • The exact pinned README states MIT for code and CC BY-SA 4.0 for content. Human review is required for attribution and ShareAlike obligations on compiled and derived data.
  • Content is declared CC BY-SA 4.0. Attribution and ShareAlike treatment of the compiled frequency evidence and derived tiers require human approval.
  • Dialect or scope: unspecified

FrequencyWords content derived from OpenSubtitles, licensed CC BY-SA 4.0

offensive-term-policy 2026.07-review-pending.2

content_filter; search_results; result_counts.

Licence: Reuse terms not yet approved. Legal review: pending. Redistribution review: pending.

  • This bounded policy list cannot determine whether every usage is offensive or benign.

The Word Index first-party content-safety exclusion policy

Citations

  1. The Word Index pinned lexical source manifest, The Word Index, version 2026.07-lexical.6 (dataset).
  2. FrequencyWords English 2018, FrequencyWords, version git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018 (dataset).
  3. The Word Index content-safety exclusion policy, The Word Index, version 2026.07-review-pending.2 (policy).

Data and reproducibility

Source-rights review is pending. The available file contains aggregate computed measures, not a redistributed spelling list. Availability does not grant broader reuse rights; consult the source and licence register before reuse.

Download the canonical aggregate CSV

View the canonical JSON study bundle.

Reproduce this version from a clean checkout:

make studies_reproduce STUDY=anagram-density