SAMPLE FRAME — AI crawler access and llms.txt adoption, collected 11 August 2026 Visibility100x. Published with the results under CC BY 4.0. https://visibility100x.com/research/ai-crawler-blocking/ There are two frames. Neither was chosen by us domain by domain, which is the only version of a sample that survives somebody asking how it was picked. FRAME 1 — "the web at large" --------------------------------------------------------------------------- Source: Tranco daily list, list ID 645KX, generated 11 August 2026. https://tranco-list.eu/list/645KX A 30 day Dowdall combination of the Chrome UX Report, Farsight, Majestic, Cloudflare Radar and Cisco Umbrella rankings, pay level domains only. Tranco is used because it is reproducible and dated; any single provider's top list is noisier and none of them publish the combination method. Selection: the top 1,500 rows of that list, in rank order, no other filter applied at selection time. Population: of those 1,500, the 1,048 that returned an HTML homepage over HTTPS to anybody. The 452 excluded are in the CSV with the reason. They are overwhelmingly infrastructure rather than websites: nameserver, CDN and ad delivery domains such as gtld-servers.net, akamaiedge.net and domaincontrol.com, which have no homepage, no content and therefore nothing for a crawler to be allowed or refused. Leaving them in would have deflated every rate on the page by roughly 30%. Known bias: Tranco ranks by global traffic and DNS resolution, so the frame is international, skews heavily technical, and over represents platforms, developer tools and infrastructure relative to ordinary businesses. It is not a sample of US companies, of any one industry, or of sites that sell anything. Read every figure as "the most visited domains on the web", not "typical websites". FRAME 2 — "the specialists" --------------------------------------------------------------------------- Source: The identical 98 domain frame published with study 1 on 6 August 2026, reused unchanged so the two studies are comparable. https://visibility100x.com/research/geo-agency-audit/ Every domain in it was named as a listed entry on a page that ranks for "best generative engine optimization agencies", "top answer engine optimization agencies" or a close variant. It is the set of agencies a buyer actually encounters. Population: all 98. Every one returned an HTML homepage. Excluded on purpose: visibility100x.com. We do not score ourselves inside our own sample. Our own result is stated separately on the page. WHAT WAS FETCHED, PER DOMAIN --------------------------------------------------------------------------- https:///robots.txt parsed with the same parser as our public AI crawler checker, RFC 9309 group selection and longest match evaluation https:///llms.txt accepted only if the body is not HTML and contains a markdown heading or link https:/// homepage, for the text a reader without JavaScript would see Each of the three was fetched twice: once declaring a research crawler, once with a Chrome user agent, from the same machine and the same address within the same hours. The comparison between those two passes is a finding in itself and is reported on the page. Timeout 15 seconds, 12 requests in flight across the whole sample, redirects followed. No cached copies, no third party database, no rendering engine.