Domain Yoga

Domain name availability benchmark — dataset

Per-name registry availability for 50,060 candidate domains across 10 TLDs, measured 2026-08-05.

Download

Results

Measured 50,060 checks on 2026-08-05 against the corpus frozen 2026-07-31. Registry RDAP resolved via IANA bootstrap, except .io and .co (absent from the bootstrap, resolved from a list maintained in this repository); 404=available, 200=registered, else indeterminate. A resolve pass re-checked 429-throttled names at 2s pacing; see resolvePasses.
Name type TLD Available Registered Unresolved Available %
compounds .ai 1250 0 0 100%
compounds .app 1248 2 0 99.8%
compounds .co 1249 1 0 99.9%
compounds .com 1215 35 0 97.2%
compounds .dev 1250 0 0 100%
compounds .io 1250 0 0 100%
compounds .net 1246 4 0 99.7%
compounds .org 1247 3 0 99.8%
compounds .tech 1250 0 0 100%
compounds .xyz 1245 5 0 99.6%
dictionary-words .ai 792 458 0 63.4%
dictionary-words .app 840 410 0 67.2%
dictionary-words .co 783 467 0 62.6%
dictionary-words .com 176 1074 0 14.1%
dictionary-words .dev 904 346 0 72.3%
dictionary-words .io 799 451 0 63.9%
dictionary-words .net 593 657 0 47.4%
dictionary-words .org 635 615 0 50.8%
dictionary-words .tech 1063 187 0 85%
dictionary-words .xyz 730 520 0 58.4%
invented .ai 1158 92 0 92.6%
invented .app 1140 110 0 91.2%
invented .co 1073 177 0 85.8%
invented .com 116 1134 0 9.3%
invented .dev 1163 87 0 93%
invented .io 1139 111 0 91.1%
invented .net 931 319 0 74.5%
invented .org 942 308 0 75.4%
invented .tech 1197 53 0 95.8%
invented .xyz 1046 204 0 83.7%
patterns .ai 1231 25 0 98%
patterns .app 1233 23 0 98.2%
patterns .co 1223 33 0 97.4%
patterns .com 1053 203 0 83.8%
patterns .dev 1248 8 0 99.4%
patterns .io 1227 29 0 97.7%
patterns .net 1224 32 0 97.5%
patterns .org 1223 33 0 97.4%
patterns .tech 1244 12 0 99%
patterns .xyz 1201 55 0 95.6%

Read this table cell by cell, not as a total. Three of the four strata are constructed rather than drawn from names people actually want, and most of the TLDs measured are not .com — so any figure pooled across the whole table describes this corpus's composition, not the domain market. What the fully crossed design buys is that the cells are comparable to each other: every label is checked against every TLD, so a difference between two cells is a difference in the TLD or the stratum, never in the sample.

One cell inverts the pattern. In .com, invented labels are registered more often than dictionary words — the reverse of the remaining 9 TLDs here, where invented names are far freer. Read it as saturation: demand has moved past the dictionary and into coined strings.

Slope chart of domain name availability across 10 extensions, measured 2026-08-05. Invented five-letter names are more available than dictionary words on 9 of 10 extensions; on .com only 9.3% of invented names are available against 14.1% of dictionary words.
.com is the only extension where invented names are scarcer than dictionary words — 9.3% against 14.1% available, measured 2026-08-05.

Method

Registry RDAP resolved via IANA bootstrap, except .io and .co (absent from the bootstrap, resolved from a list maintained in this repository); 404=available, 200=registered, else indeterminate. A resolve pass re-checked 429-throttled names at 2s pacing; see resolvePasses.

Corpus frozen 2026-07-31 (version 1); measured 2026-08-05. 5,006 names across 10 TLDs — 50,060 checks in total, 0 unresolved.

This capture spans 10 TLDs across 7 registries. Each check resolves the name's registry RDAP record, ordinarily via IANA's bootstrap: a 404 means available, a 200 means registered, anything else is recorded as indeterminate rather than guessed. 2 TLDs are absent from IANA's bootstrap and resolve instead from a list maintained in this repository: .io, .co.

Before any candidate name was measured, the capture script probed one known-registered domain per TLD and would have aborted the run had any come back other than registered — the defence against a wrong-but-real RDAP endpoint, which answers a registered name with a well-formed 404 indistinguishable from a genuine free-name response, with nothing in the resulting data to reveal the failure afterward.

The corpus

The same 5,006 names are measured every run, and that is what makes runs comparable. A claim like "availability on .com is down since last month" is only honest if the names didn't change in between. The TLD set is free to widen, though: each stratum × TLD observation stands on its own, so measuring more TLDs in a later run doesn't disturb the comparisons that already exist.

Four strata. Three draw from web2 — Webster's Second International, 1934, chosen because it is public domain: 1,250 dictionary-words are single web2 words, 1,250 compounds are two web2 words concatenated, and 1,256 patterns apply a prefix or suffix (get-, try-, -hq and similar) to a web2 word. The fourth stratum, 1,250 invented names, is pronounceable CVCVC strings generated independently of any wordlist.

The corpus is not published as a standalone file — it lives in a private source repository, version 1, frozen 2026-07-31. That is not the same as hidden: every one of its 5,006 labels, its stratum, and — for patterns — the specific pattern it is built from is recoverable from the NDJSON above, because each measurement record names the exact candidate that was checked.

Known bias

web2 is a dictionary, not a frequency list. The real-words and compounds strata over-represent archaic and obscure vocabulary, which is more available than common words. Read those two strata as an optimistic ceiling. Do not label them 'common English words'. A frequency-ranked list would fix this and needs sourcing. Separately: 'available' here means 'no registration record', not 'registrable'. A registry-reserved name has no registration record either, so it reads the same as a genuinely free one — 0 of the 15,018 published 2026-08-02 records carry the 'reserved' verdict, even though 'reserved' sits in the published enum in a way that could read as implying this method can detect it. RDAP has no standard signal for the distinction: RFC 9083 defines no reserved-name field and ICANN's gTLD RDAP profile requires none. Verified 2026-08-05 against rdap.nic.tv: the reserved afrostream.tv and an unregistered .tv name return byte-identical responses (404, errorCode 404, title 'Not found', description 'No data found'), while WHOIS for the same reserved name returns 'Reserved Domain Name' and the free name does not. A registry could differentiate via its own errorCode or title text; nothing requires it and none measured here does. So this is 'no standard signal, and none observed' — not proof RDAP can never distinguish the two. 'reserved' remains valid in this schema for a corroborating source (WHOIS or zone) that can make the call; an RDAP-only run cannot produce it.

The passage above is reproduced verbatim from the frozen corpus, including its own open item about needing a frequency-ranked list — we haven't rewritten it, because doing that would mean editing data inside a frozen, versioned file after the fact. It also says "real-words"; the corpus and this table call that same stratum dictionary-words.

The bias isn't confined to dictionary-words and compounds: every one of the 157 distinct bases the patterns stratum applies get-/try-/-hq-style prefixes and suffixes to is itself a dictionary-words word (getaalii, tryacetone, adonitehq — not typical startup vocabulary). Patterns rows are comparable with each other — one pattern scoring a few points ahead of another on the same TLD is a real finding — but their absolute availability rates inherit the same optimistic ceiling, and shouldn't be read as evidence that "a normal word with get-/try-/-hq" is generally this available.

2,079 of 50,060 checks (4.2%) were still unresolved after the capture script's inline retries, and were re-checked across 1 paced pass recorded in resolvePasses, leaving 0 unresolved. Throttling is a property of the registries, not of the names: it concentrates in whichever registries rate-limit, and the per-check attempts and rdapHttpStatus fields in the NDJSON above show exactly which.

Licence

The measurements and this compilation are released under CC BY 4.0. Please credit: Domain Yoga — https://domain.yoga/data/availability. The licence covers our measurements and their compilation, not the underlying wordlist, which is Webster’s Second (1934) and in the public domain.

Cite this dataset

This dataset is archived with a DOI, which resolves to its current version: 10.5281/zenodo.21814388 . Prefer it to this page’s address — the DOI survives a reorganisation of this site, and always leads to the newest capture.

Domain Yoga. (2026). Domain name availability benchmark [Data set]. Captured 2026-08-05, measured through 2026-08-06 via resolve passes. Zenodo. https://doi.org/10.5281/zenodo.21814388. Licensed under CC BY 4.0.