Word-ladder connectivity records and classic doublets
Published for direct use; search approval pending. Its mechanically validated result is available, but the page remains noindex and outside sitemaps while human editorial and source-rights review is incomplete.
Result
Across lengths 3-7, 51,741 eligible records form 101,904 one-letter substitution links; the longest certified shortest ladder found spans 52 steps, and 10 of 10 classic doublet pairs have a ladder in this build.
- Analysis population
- 51,741
- Eligible ladder records
- 51,741
- Eligible records with no substitution neighbor
- 22.91%
- Longest certified shortest ladder found
- 52
- Classic doublet pairs solvable in this build
- 10
Across lengths 3-7, 51,741 eligible records form 101,904 one-letter substitution links; the longest certified shortest ladder found spans 52 steps, and 10 of 10 classic doublet pairs have a ladder in this build.
How the records are certified
Longest-ladder values come from deterministic farthest-vertex sweeps: each is an exact shortest path between two eligible records, so it is a certified lower bound and the true maximum may be longer. Vertex eligibility uses build-specific ranking tiers and the product safety policy; it does not establish universal English usage or acceptance in any game.
Eligible records and islands by length
Exact per-length counts appear in the adjacent table.
Population and exclusions
| Population stage | Records (records) |
|---|---|
| Raw lexical ledger | 394595 |
| Excluded by product policy | 150 |
| Safe normalized-spelling records | 394445 |
| Records meeting this study's inclusion criteria | 51741 |
Headline ladder records
| Measure | Count (items) | Share of eligible records (percent) | Ladder length (steps) (steps) |
|---|---|---|---|
| Eligible records across lengths 3-7 | 51741 | — | — |
| Substitution links | 101904 | — | — |
| Island records with no substitution neighbor | 11855 | 22.91 | — |
| Classic doublet pairs with a ladder in this build | 10 | — | — |
| Longest certified shortest ladder found | — | — | 52 |
Connectivity by spelling length
| Letters | Eligible records (records) | Substitution links (edges) | Island records (records) | Components (components) | Largest component (records) | Largest-component share (percent) | Longest certified ladder found (steps) |
|---|---|---|---|---|---|---|---|
| 3 | 970 | 7767 | 5 | 6 | 965 | 99.48 | 9 |
| 4 | 3880 | 20740 | 66 | 78 | 3785 | 97.55 | 17 |
| 5 | 8604 | 24635 | 743 | 913 | 7364 | 85.59 | 30 |
| 6 | 15204 | 24721 | 3259 | 4374 | 8385 | 55.15 | 39 |
| 7 | 23083 | 24041 | 7782 | 10277 | 7634 | 33.07 | 52 |
Classic doublets checked against this build
| Doublet | Letters | Eligible endpoints (records) | Shortest ladder (steps) (steps) |
|---|---|---|---|
| ape to man | 3 | 2 | 5 |
| eel to pie | 3 | 2 | 4 |
| pig to sty | 3 | 2 | 5 |
| cold to warm | 4 | 2 | 4 |
| four to five | 4 | 2 | 6 |
| head to tail | 4 | 2 | 4 |
| flour to bread | 5 | 2 | 6 |
| tears to smile | 5 | 2 | 6 |
| wheat to bread | 5 | 2 | 6 |
| winter to summer | 6 | 2 | 6 |
Methodology
Research question: How connected are this build's one-letter-substitution ladder graphs at each length, how long are the longest certified shortest ladders, and which classic doublet pairs remain solvable here?
Why it is useful: The records show whether a ladder is possible between two spellings in this build and explain why a pair may lack a path under the declared source and eligibility rules.
Included
word-ladder-records.vertex-eligibility: Include safe normalized spellings of lengths three through seven with ranking tier three or better as graph vertices.word-ladder-records.edge-rule: Connect two vertices of the same length exactly when they differ by one letter substitution.
Excluded
word-ladder-records.product-policy: Exclude spellings matched by the exact first-party content-safety policy before building the graphs.word-ladder-records.outside-length-window: Exclude spellings shorter than three or longer than seven letters from the graphs.
Transformations
word-ladder-records.graph-build-v1: Bucket vertices by single-blank hole patterns to enumerate substitution edges, then derive components, islands, and certified path records with deterministic breadth-first searches.
Calculations
research.count-v1- Count vertices, edges, islands, components, or solvable doublet pairs in each declared group. Formula:
count(records). research.share-v1- Calculate the exact share of eligible records inside a declared connectivity group. Formula:
100 * group_count / eligible_records. research.maximum-v1- Take the largest connected-component size within a declared length graph. Formula:
max(component_size). research.shortest-path-v1- Find the exact fewest substitution steps between two named eligible endpoints. Formula:
min_steps(breadth_first_search(start, target)). research.longest-shortest-path-found-v1- Take the longest certified breadth-first distance found by deterministic farthest-vertex sweeps over every component. Formula:
max(certified_bfs_distance over sweeps).
Limitations
- Vertex eligibility uses build-specific ranking bands and the product safety policy, so connectivity describes this exact build only.
- The longest-ladder records are certified lower bounds from deterministic sweeps, not exhaustive all-pairs searches.
- Classic doublet endpoints are quoted from a public-domain 1879 puzzle; historical solutions may use spellings this build excludes.
- The aggregate download contains counts, step records, and the quoted doublet endpoints only; it does not redistribute word lists.
Versions, sources, and review
dwyl-english-words git:8179fe68775df3f553ef19520db065228e65d1d3
source-membership; surface-spelling-as-listed.
Licence: Unlicense notice; upstream dataset rights unresolved. Legal review: pending. Redistribution review: pending.
- The pinned repository licence is the Unlicense, but the pinned README says copyright in the extracted source list remains with Infochimps. Human legal review of the upstream rights chain is required.
- Do not assert a redistribution right until the Infochimps-to-dwyl rights chain has been reviewed by a human.
- Dialect or scope: unspecified
dwyl/english-words; the repository carries an Unlicense notice, while its pinned README attributes upstream copyright to Infochimps
enable artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7
source-membership; surface-spelling-as-listed.
Licence: Licence terms not stated by the source. Legal review: pending. Redistribution review: pending.
- The downloaded artifact is content-pinned by SHA-256, but it contains no licence notice and the configured source page does not state reuse terms.
- The artifact is reproducibly pinned but no affirmative reuse grant has been located; redistribution approval remains pending.
- Dialect or scope: North-American-oriented; exact edition metadata unavailable
ENABLE (Enhanced North American Benchmark Lexicon); exact reuse terms are pending legal review
google-10000-english git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8
source-membership; source-order.
Licence: No explicit licence grant in pinned repository metadata. Legal review: pending. Redistribution review: pending.
- The exact pinned LICENSE.md describes provenance but contains no explicit public-domain dedication or licence grant. Human review is required.
- The pinned repository documents provenance but does not supply an affirmative licence grant; do not describe it as public domain.
- Dialect or scope: USA no-swears variant
first20hours/google-10000-english, derived from the Google Web Trillion Word Corpus and Peter Norvig's compilation
frequencywords-en-2018 git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018
source-token-observation; source-token-count.
Licence: CC BY-SA 4.0 (content). Legal review: pending. Redistribution review: pending.
- The exact pinned README states MIT for code and CC BY-SA 4.0 for content. Human review is required for attribution and ShareAlike obligations on compiled and derived data.
- Content is declared CC BY-SA 4.0. Attribution and ShareAlike treatment of the compiled frequency evidence and derived tiers require human approval.
- Dialect or scope: unspecified
FrequencyWords content derived from OpenSubtitles, licensed CC BY-SA 4.0
offensive-term-policy 2026.07-review-pending.2
content_filter; search_results; result_counts.
Licence: Reuse terms not yet approved. Legal review: pending. Redistribution review: pending.
- This bounded policy list cannot determine whether every usage is offensive or benign.
The Word Index first-party content-safety exclusion policy
Citations
- The Word Index pinned lexical source manifest, The Word Index, version 2026.07-lexical.6 (dataset).
- FrequencyWords English 2018, FrequencyWords, version git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018 (dataset).
- The Word Index content-safety exclusion policy, The Word Index, version 2026.07-review-pending.2 (policy).
- Doublets: a verbal puzzle, Vanity Fair (1879), version 1879 (method).
Data and reproducibility
Source-rights review is pending. The available file contains aggregate computed measures, not a redistributed spelling list. Availability does not grant broader reuse rights; consult the source and licence register before reuse.
Download the canonical aggregate CSV
View the canonical JSON study bundle.
Reproduce this version from a clean checkout:
make studies_reproduce STUDY=word-ladder-records