The Delve Index: a self-audit of our machine-written definition corpus

Published for direct use; search approval pending. Its mechanically validated result is available, but the page remains noindex and outside sitemaps while human editorial and source-rights review is incomplete.

Result

First-party glosses document 59.98% of safe records; the 22 declared style markers occur 846 times in 2,973,912 gloss tokens.

Analysis population
394,445
Safe records documented by first-party glosses
59.98%
Combined occurrences of the declared style markers
846
Most-used rarer-band gloss token
participle

Methodology first: this corpus audits itself

The definition glosses are first-party, machine-generated text governed by this product's review policy; inclusion in the build is not word-level editorial approval. This study counts exact lowercase a-z token runs in those glosses and joins each token to its own compiled record, so every rate describes the corpus's writing style, not English.

The style-marker list is a fixed editorial heuristic recorded in version control. Appearing on it is not evidence about any individual gloss.

First-party glosses document 59.98% of safe records; the 22 declared style markers occur 846 times in 2,973,912 gloss tokens.

What the audit cannot claim

Token counting ignores senses, phrases, and grammar; a marker can be the only correct word in context. The comparison population is this build's safe compiled lexicon, so shares are build-specific and do not measure general English usage or the style of any other text.

Gloss token share versus safe record share by tier

Exact shares and denominators appear in the adjacent table.

Read the accessible data table.

Population and exclusions

The exact population used by this version of the analysis.
Population stage Records (records)
Raw lexical ledger 394595
Excluded by product policy 150
Safe normalized-spelling records 394445
Records meeting this study's inclusion criteria 394445

What was counted

Corpus sizes are exact counts over the sealed first-party glosses.
Metric Observations (observations)
Safe records with at least one first-party gloss 236598
First-party definition glosses 271328
Gloss token occurrences (a-z runs) 2973912
Distinct gloss tokens 103776
Distinct gloss tokens that are safe compiled records 99627

Gloss coverage

Coverage describes this build's first-party corpus only.
Coverage class Records (records) Share (percent)
Records with first-party glosses 236598 59.98
Records without first-party glosses 157847 40.02

Where gloss vocabulary sits in the lexicon

Token shares and record shares use different exact denominators, both named in the dataset.
Cohort Gloss token occurrences (tokens) Share of gloss tokens (percent) Safe records at tier (records) Share of safe records (percent)
Tier 1 safe records 2228560 74.94 9943 2.52
Tier 2 safe records 430085 14.46 15100 3.83
Tier 3 safe records 251707 8.46 147637 37.43
Tier 4 safe records 29018 0.98 1744 0.44
Tier 5 safe records 27980 0.94 220021 55.78
Tokens that are not safe compiled records 6562 0.22

Most-used tier 3-5 vocabulary in the glosses

A computed exemplar limited to 25 rows; policy-excluded spellings can never appear.
Gloss token Tier of the matching record Gloss occurrences (tokens) Documented headwords using it (records) Occurrences per 10,000 tokens (per-10k-tokens)
participle 3 12562 12430 42.24
s 5 5446 5103 18.31
is 3 4706 4555 15.82
genus 3 3720 3689 12.51
british 4 3258 3156 10.96
american 4 1804 1780 6.07
scottish 4 1770 1693 5.95
dialectal 3 1738 1715 5.84
superlative 3 1723 1720 5.79
it 3 1565 1525 5.26
latin 4 1044 1029 3.51
asia 4 821 814 2.76
america 4 805 798 2.71
abbreviation 3 801 717 2.69
spanish 4 781 759 2.63
th 4 717 562 2.41
italian 4 706 666 2.37
african 4 603 590 2.03
africa 4 596 587 2.00
first 3 535 499 1.80
australian 4 479 465 1.61
ornamental 3 464 457 1.56
chiefly 3 438 432 1.47
notably 3 437 437 1.47
asian 4 435 429 1.46

The declared style-marker list, in full

Zero counts are published so the audit cannot cherry-pick.
Marker Tier of the matching record Gloss occurrences (tokens) Distinct documented headwords using it (records) Occurrences per 10,000 tokens (per-10k-tokens)
boast 2 29 26 0.10
crucial 1 12 12 0.04
delve 2 6 6 0.02
emphasize 2 50 50 0.17
encompass 3 14 14 0.05
evoke 2 21 20 0.07
facilitate 2 24 24 0.08
foster 1 15 13 0.05
intricate 2 56 54 0.19
leverage 1 7 6 0.02
meticulous 2 13 13 0.04
multifaceted 3 0 0 0.00
myriad 2 1 1 0.00
notably 3 437 437 1.47
nuanced 3 1 1 0.00
pivotal 2 2 2 0.01
realm 1 101 100 0.34
robust 2 25 25 0.08
showcase 2 4 3 0.01
tapestry 2 15 13 0.05
underscore 3 7 5 0.02
vibrant 2 6 5 0.02
All declared markers combined 846 826 2.84

Methodology

Research question: Which tier 3-5 compiled tokens occur most often in the first-party definition corpus, and how often do declared machine-prose style markers appear?

Why it is useful: Documents, with exact counts, which tier 3-5 tokens recur most often in this site's machine-written definitions and how frequently the declared style markers appear.

Included

  • delve-index.first-party-glosses: Include every first-party definition gloss attached to a safe compiled record in this release.
  • delve-index.safe-lexicon-join: Join each distinct gloss token to the safe compiled record with the same normalized spelling, where one exists.

Excluded

  • delve-index.product-policy: Exclude spellings matched by the exact first-party content-safety policy; excluded spellings can never appear as joined tokens.
  • delve-index.non-alphabetic-runs: Tokens are exact lowercase a-z runs; digits, punctuation, and whitespace delimit tokens and are never counted.

Transformations

  1. delve-index.tokenize-v1: Casefold every first-party gloss and extract exact lowercase a-z token runs.
  2. delve-index.join-v1: Join each distinct token to the safe compiled record with the same spelling and attribute occurrences to that record's tier.
  3. delve-index.audit-v1: Count occurrences of the fixed declared style-marker list, count the distinct documented headwords containing any marker, and rank tier 3-5 tokens by exact occurrences, limited to 25 rows.

Calculations

research.count-v1
Count records, glosses, and exact lowercase a-z token runs. Formula: count(observations).
research.share-v1
Calculate a cohort's exact share of its named denominator. Formula: 100 * subset_count / cohort_count.
delve.rate-per-10k-v1
Normalize exact token occurrences per ten thousand gloss tokens. Formula: 10000 * token_occurrences / total_gloss_tokens.
delve.style-marker-count-v1
Count exact occurrences and distinct documented headwords for a fixed version-controlled marker list associated with machine-written explanatory prose, publishing zero counts; the combined headword total uses a set union. Formula: count(token occurrences where token in declared marker list); count(distinct documented headwords containing one or more declared markers).
delve.token-leader-v1
Select the most-frequent rarer-band gloss token, breaking equal counts by ascending token. Formula: argmax(gloss_occurrences, tie_break=token_ascending).

Limitations

  • The glosses are first-party machine-generated text; every rate describes that corpus's writing style, not English usage.
  • Token counting ignores senses, multi-word phrases, and grammar; a marker can be the only correct word in its context.
  • Splitting on non-letters fragments possessives and abbreviations, so single-letter tokens such as 's' occur.
  • The style-marker list is a fixed editorial heuristic recorded in version control.
  • Tier joins are build-specific ranking bands, not universal usage claims.

Versions, sources, and review

dwyl-english-words git:8179fe68775df3f553ef19520db065228e65d1d3

source-membership; surface-spelling-as-listed.

Licence: Unlicense notice; upstream dataset rights unresolved. Legal review: pending. Redistribution review: pending.

  • The pinned repository licence is the Unlicense, but the pinned README says copyright in the extracted source list remains with Infochimps. Human legal review of the upstream rights chain is required.
  • Do not assert a redistribution right until the Infochimps-to-dwyl rights chain has been reviewed by a human.
  • Dialect or scope: unspecified

dwyl/english-words; the repository carries an Unlicense notice, while its pinned README attributes upstream copyright to Infochimps

enable artifact-sha256:d3fbe8485022088fcf527edcde2fbdc18b4bbc141ac58123c9adb462e086eaf7

source-membership; surface-spelling-as-listed.

Licence: Licence terms not stated by the source. Legal review: pending. Redistribution review: pending.

  • The downloaded artifact is content-pinned by SHA-256, but it contains no licence notice and the configured source page does not state reuse terms.
  • The artifact is reproducibly pinned but no affirmative reuse grant has been located; redistribution approval remains pending.
  • Dialect or scope: North-American-oriented; exact edition metadata unavailable

ENABLE (Enhanced North American Benchmark Lexicon); exact reuse terms are pending legal review

google-10000-english git:bdf4c221bc120b0b7f6c3f1eff1cc1abb975f8d8

source-membership; source-order.

Licence: No explicit licence grant in pinned repository metadata. Legal review: pending. Redistribution review: pending.

  • The exact pinned LICENSE.md describes provenance but contains no explicit public-domain dedication or licence grant. Human review is required.
  • The pinned repository documents provenance but does not supply an affirmative licence grant; do not describe it as public domain.
  • Dialect or scope: USA no-swears variant

first20hours/google-10000-english, derived from the Google Web Trillion Word Corpus and Peter Norvig's compilation

frequencywords-en-2018 git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018

source-token-observation; source-token-count.

Licence: CC BY-SA 4.0 (content). Legal review: pending. Redistribution review: pending.

  • The exact pinned README states MIT for code and CC BY-SA 4.0 for content. Human review is required for attribution and ShareAlike obligations on compiled and derived data.
  • Content is declared CC BY-SA 4.0. Attribution and ShareAlike treatment of the compiled frequency evidence and derived tiers require human approval.
  • Dialect or scope: unspecified

FrequencyWords content derived from OpenSubtitles, licensed CC BY-SA 4.0

offensive-term-policy 2026.07-review-pending.2

content_filter; search_results; result_counts.

Licence: Reuse terms not yet approved. Legal review: pending. Redistribution review: pending.

  • This bounded policy list cannot determine whether every usage is offensive or benign.

The Word Index first-party content-safety exclusion policy

Citations

  1. The Word Index first-party definition corpus, The Word Index, version 2026.07-lexical.6 (dataset).
  2. The Word Index pinned lexical source manifest, The Word Index, version 2026.07-lexical.6 (dataset).
  3. FrequencyWords English 2018, FrequencyWords, version git:525f9b560de45753a5ea01069454e72e9aa541c6; corpus-release:2018 (dataset).
  4. The Word Index content-safety exclusion policy, The Word Index, version 2026.07-review-pending.2 (policy).

Data and reproducibility

Source-rights review is pending. The available file contains aggregate computed measures, not a redistributed spelling list. Availability does not grant broader reuse rights; consult the source and licence register before reuse.

Download the canonical aggregate CSV

View the canonical JSON study bundle.

Reproduce this version from a clean checkout:

make studies_reproduce STUDY=delve-index