Domain name availability benchmark — dataset
Per-name registry availability for 50,060 candidate domains across 10 TLDs, measured 2026-08-05.
Download
- Per-name measurements (NDJSON) — one record per candidate domain checked: label, TLD, stratum, pattern, verdict, and sources.
- Aggregate + provenance (JSON) — capture date, method, registry endpoints, licence, credit line, known bias, and the by-stratum/by-pattern rollups behind the table below.
Results
| Name type | TLD | Available | Registered | Unresolved | Available % |
|---|---|---|---|---|---|
| compounds | .ai | 1250 | 0 | 0 | 100% |
| compounds | .app | 1248 | 2 | 0 | 99.8% |
| compounds | .co | 1249 | 1 | 0 | 99.9% |
| compounds | .com | 1215 | 35 | 0 | 97.2% |
| compounds | .dev | 1250 | 0 | 0 | 100% |
| compounds | .io | 1250 | 0 | 0 | 100% |
| compounds | .net | 1246 | 4 | 0 | 99.7% |
| compounds | .org | 1247 | 3 | 0 | 99.8% |
| compounds | .tech | 1250 | 0 | 0 | 100% |
| compounds | .xyz | 1245 | 5 | 0 | 99.6% |
| dictionary-words | .ai | 792 | 458 | 0 | 63.4% |
| dictionary-words | .app | 840 | 410 | 0 | 67.2% |
| dictionary-words | .co | 783 | 467 | 0 | 62.6% |
| dictionary-words | .com | 176 | 1074 | 0 | 14.1% |
| dictionary-words | .dev | 904 | 346 | 0 | 72.3% |
| dictionary-words | .io | 799 | 451 | 0 | 63.9% |
| dictionary-words | .net | 593 | 657 | 0 | 47.4% |
| dictionary-words | .org | 635 | 615 | 0 | 50.8% |
| dictionary-words | .tech | 1063 | 187 | 0 | 85% |
| dictionary-words | .xyz | 730 | 520 | 0 | 58.4% |
| invented | .ai | 1158 | 92 | 0 | 92.6% |
| invented | .app | 1140 | 110 | 0 | 91.2% |
| invented | .co | 1073 | 177 | 0 | 85.8% |
| invented | .com | 116 | 1134 | 0 | 9.3% |
| invented | .dev | 1163 | 87 | 0 | 93% |
| invented | .io | 1139 | 111 | 0 | 91.1% |
| invented | .net | 931 | 319 | 0 | 74.5% |
| invented | .org | 942 | 308 | 0 | 75.4% |
| invented | .tech | 1197 | 53 | 0 | 95.8% |
| invented | .xyz | 1046 | 204 | 0 | 83.7% |
| patterns | .ai | 1231 | 25 | 0 | 98% |
| patterns | .app | 1233 | 23 | 0 | 98.2% |
| patterns | .co | 1223 | 33 | 0 | 97.4% |
| patterns | .com | 1053 | 203 | 0 | 83.8% |
| patterns | .dev | 1248 | 8 | 0 | 99.4% |
| patterns | .io | 1227 | 29 | 0 | 97.7% |
| patterns | .net | 1224 | 32 | 0 | 97.5% |
| patterns | .org | 1223 | 33 | 0 | 97.4% |
| patterns | .tech | 1244 | 12 | 0 | 99% |
| patterns | .xyz | 1201 | 55 | 0 | 95.6% |
Read this table cell by cell, not as a total. Three of the four strata are constructed
rather than drawn from names people actually want, and most of the TLDs measured are not
.com — so any figure pooled across the whole table describes this corpus's
composition, not the domain market. What the fully crossed design buys is that the cells
are comparable to each other: every label is checked against every TLD, so a
difference between two cells is a difference in the TLD or the stratum, never in the sample.
One cell inverts the pattern.
In .com, invented labels are
registered more often than dictionary words — the reverse of the remaining 9 TLDs here, where invented names are far
freer. Read it as saturation: demand has moved past the dictionary and into coined
strings.
Method
Registry RDAP resolved via IANA bootstrap, except .io and .co (absent from the bootstrap, resolved from a list maintained in this repository); 404=available, 200=registered, else indeterminate. A resolve pass re-checked 429-throttled names at 2s pacing; see resolvePasses.
Corpus frozen 2026-07-31 (version 1); measured 2026-08-05. 5,006 names across 10 TLDs — 50,060 checks in total, 0 unresolved.
This capture spans 10 TLDs across 7 registries. Each check resolves the name's registry RDAP record, ordinarily via IANA's bootstrap: a 404 means available, a 200 means registered, anything else is recorded as indeterminate rather than guessed. 2 TLDs are absent from IANA's bootstrap and resolve instead from a list maintained in this repository: .io, .co.
Before any candidate name was measured, the capture script probed one known-registered domain per TLD and would have aborted the run had any come back other than registered — the defence against a wrong-but-real RDAP endpoint, which answers a registered name with a well-formed 404 indistinguishable from a genuine free-name response, with nothing in the resulting data to reveal the failure afterward.
The corpus
The same 5,006 names are measured every run, and that is what makes runs comparable. A
claim like "availability on .com is down since last month" is only honest if the
names didn't change in between. The TLD set is free to widen, though: each stratum × TLD
observation stands on its own, so measuring more TLDs in a later run doesn't disturb the
comparisons that already exist.
Four strata. Three draw from web2 — Webster's Second International, 1934, chosen
because it is public domain: 1,250 dictionary-words are single web2 words, 1,250 compounds are two web2 words concatenated, and 1,256 patterns apply a prefix or suffix (get-, try-, -hq and similar) to a web2 word.
The fourth stratum, 1,250 invented names, is pronounceable CVCVC
strings generated independently of any wordlist.
The corpus is not published as a standalone file — it lives in a private source repository, version 1, frozen 2026-07-31. That is not the same as hidden: every one of its 5,006 labels, its stratum, and — for patterns — the specific pattern it is built from is recoverable from the NDJSON above, because each measurement record names the exact candidate that was checked.
Known bias
web2 is a dictionary, not a frequency list. The real-words and compounds strata over-represent archaic and obscure vocabulary, which is more available than common words. Read those two strata as an optimistic ceiling. Do not label them 'common English words'. A frequency-ranked list would fix this and needs sourcing. Separately: 'available' here means 'no registration record', not 'registrable'. A registry-reserved name has no registration record either, so it reads the same as a genuinely free one — 0 of the 15,018 published 2026-08-02 records carry the 'reserved' verdict, even though 'reserved' sits in the published enum in a way that could read as implying this method can detect it. RDAP has no standard signal for the distinction: RFC 9083 defines no reserved-name field and ICANN's gTLD RDAP profile requires none. Verified 2026-08-05 against rdap.nic.tv: the reserved afrostream.tv and an unregistered .tv name return byte-identical responses (404, errorCode 404, title 'Not found', description 'No data found'), while WHOIS for the same reserved name returns 'Reserved Domain Name' and the free name does not. A registry could differentiate via its own errorCode or title text; nothing requires it and none measured here does. So this is 'no standard signal, and none observed' — not proof RDAP can never distinguish the two. 'reserved' remains valid in this schema for a corroborating source (WHOIS or zone) that can make the call; an RDAP-only run cannot produce it.
The passage above is reproduced verbatim from the frozen corpus, including its own open
item about needing a frequency-ranked list — we haven't rewritten it, because doing that
would mean editing data inside a frozen, versioned file after the fact. It also says
"real-words"; the corpus and this table call that same stratum dictionary-words.
The bias isn't confined to dictionary-words and compounds: every
one of the 157 distinct bases the patterns stratum applies get-/try-/-hq-style
prefixes and suffixes to is itself a dictionary-words word (getaalii, tryacetone,
adonitehq — not typical startup vocabulary). Patterns rows are comparable with each other —
one pattern scoring a few points ahead of another on the same TLD is a real finding — but
their absolute availability rates inherit the same optimistic ceiling, and shouldn't be read
as evidence that "a normal word with get-/try-/-hq" is generally this available.
2,079 of 50,060 checks
(4.2%) were still
unresolved after the capture script's inline retries, and were re-checked across 1 paced pass recorded in resolvePasses, leaving 0 unresolved.
Throttling is a property of the registries, not of the names: it concentrates in
whichever registries rate-limit, and the per-check attempts and rdapHttpStatus fields in the NDJSON above show exactly which.
Licence
The measurements and this compilation are released under CC BY 4.0. Please credit: Domain Yoga — https://domain.yoga/data/availability. The licence covers our measurements and their compilation, not the underlying wordlist, which is Webster’s Second (1934) and in the public domain.
Cite this dataset
This dataset is archived with a DOI, which resolves to its current version: 10.5281/zenodo.21814388 . Prefer it to this page’s address — the DOI survives a reorganisation of this site, and always leads to the newest capture.
Domain Yoga. (2026). Domain name availability benchmark [Data set]. Captured 2026-08-05, measured through 2026-08-06 via resolve passes. Zenodo. https://doi.org/10.5281/zenodo.21814388. Licensed under CC BY 4.0.