We audited 98 GEO agencies with our own engine
Every agency named by the listicles that rank for this category, put through the same 36 checks we sell. The full data is below, with no form to fill in.
We audited 98 agencies selling generative engine optimization, running 36 technical checks against each of their own websites. 10 refused our fetcher, so the rates below are over the 88 we could read. Most are technically sound: the median is 87 out of 100 and nothing scores below a B. But 28 of 88 publish pages carrying no date at all, and 22 name no author in the markup. Those are the two signals they sell.
The two things they most often get wrong are the two they sell
This was not the expected result. The working assumption going in was that most agency sites would be a mess. They are not: the distribution is tight and mostly good. The story is in which specific checks fail.
- No date at all
- 28 of 88
- No author
- 22 of 88
- Median score
- 87
- Measured twice
- 11 of 13
32% carry no publish date, no modified date and no visible date.
25% name nobody in the markup and carry no Person or Organization schema.
out of 100. Mean 86, range 75 to 93.
checks returned identical counts on both runs, including both headline figures. The two that moved measure the network.
Freshness and attribution are not obscure hygiene items. They are the two signals every agency in this sample tells its own clients to fix, because generative engines lean on them hardest when deciding whether a page is worth quoting: an undated page reads as undated, and an unattributed claim reads as a claim. Around a third of the category publishes undated pages and a quarter publishes unattributed ones, and sells the fix for both.
Being precise about this matters more than the number, because the obvious objection is that plenty of pages show a date a reader can see even when the markup does not carry one. That objection does not apply here. No publish date, no modified date, and no date visible to a reader either. Every one of these sites scored zero on this check rather than merely low, so this is not a case of a date being present but unmarked.
The attribution failures are equally unambiguous. No author named anywhere in the markup, and no Person or Organization schema on the page. Some of these sites do show a byline a reader can see, with nothing behind it a parser can read, so a model has nothing to attribute the page to beyond the domain name.
And four of them refuse the crawlers they sell access to
4 of the 88 block at least one major AI crawler in their own robots file, on the homepage of a business selling visibility inside AI answers. Blocking those crawlers is a legitimate choice for a publisher protecting its content. It is a harder position to hold while selling generative engine optimisation, and it is the cheapest claim here to check: open the robots file.
Two rows in the table are worth as much as any of the failures. Render gap: 0 of 88. Content depth: 0. The expectation going in was that a category selling JavaScript rendering fixes would contain a few sites with the defect. It contains none, and a study that quietly dropped its zero rows would be reporting the half of its own result that suits it.
We are not naming individual agencies in the prose, and the reason is not delicacy. A study that leads with a list of embarrassed companies gets read as an attack and shared by nobody except the people it flatters. The full CSV carries every domain and every result, so anybody who wants to check a specific claim, including any agency in it, can.
Every check that failed, and what failing it means
Counts out of the 88 sites we could fetch, not out of 98. A page that answered HTTP 403 has no content to check, and counting it as a failure on every content check is how the first pass at this data produced rates a third too high. Both collections are shown, because one run is one sample of a moving target.
| Check | 2026-08-05 | 2026-08-06 | What a failure means |
|---|---|---|---|
| Canonical target | 31 35% | 31 35% | The canonical tag points somewhere that redirects or does not resolve cleanly. |
| Freshness signals | 28 32% | 28 32% | The page carries no date a machine or a reader can find. |
| Attribution and sourcing | 22 25% | 22 25% | No author in the markup and no Person or Organization schema. |
| Page weight | 22 25% | 22 25% | The page ships more bytes than its content justifies. |
| Response time | 19 22% | 21 24% | The server took long enough that a crawler on a budget may not wait. |
| Structured data | 11 13% | 11 13% | Schema is absent or fails to parse, in which case it is silently ignored. |
| Caching headers | 10 11% | 10 11% | No usable cache policy, so repeat visitors and crawlers refetch everything. |
| AI crawler access | 4 5% | 4 5% | At least one major AI crawler is refused in robots.txt. |
| Headings | 3 3% | 3 3% | No H1, or a heading structure a parser cannot follow. |
| Image alt coverage | 3 3% | 3 3% | Most images carry no alt text. |
| Core Web Vitals | 3 3% | 8 9% | Failing Google's own thresholds for loading, interactivity or layout shift. |
| Render gap | 0 0% | 0 0% | Content appears only after JavaScript runs, so a crawler that does not run scripts receives a shell. |
| Content depth | 0 0% | 0 0% | Too little substantive content on the page to be useful as a source. |
Grades, and the category averages
Methodology
The sample. 98 domains, every one named as a listed entry on a page currently ranking for "best generative engine optimization agencies", "top answer engine optimization agencies" or a close variant. That is a deliberate frame: it is the set of agencies a buyer actually encounters, not a set we chose. The frame and the source pages it was harvested from are published as a plain text file.
What was measured. Each agency's own homepage, put through the same 36check engine behind our free audit, reading the live HTTP response. No sampling of pages within a site, no cached copies and no estimates from a database.
Concurrency 2, and this is a measurement decision rather than a courtesy. Several checks call Google PageSpeed Insights. At three or more requests in flight PSI begins throttling and returns degraded numbers, which show up as grades that are an artefact of how fast we ran the study. A score that depends on our concurrency is not a score.
Completion, and the denominator. 98 of 98 returned a result with 0 timeouts or DNS misses. 10 of those answeredHTTP 403: bot protection reacting to an unrecognised user agent rather than sites that are down, and all ten load in a browser. A blocked fetch returns no HTML, so every content check on it would record a failure it did not earn. Those 10 are excluded from every rate on this page and are in the CSV with their status. It is worth being explicit because a study that silently loses part of its sample is a study about crawlability by accident, and its headline rates are computed over a denominator it does not disclose.
Collected twice, and both runs are published. The full sample was re run on 2026-08-06 before publication, and both sets of counts are in the table above rather than only the first. A single collection is one sample of a moving target.
The reproducibility result is worth reporting on its own. Thirteen of the fifteen checks returned identical counts across the two runs, including both headline figures: freshness 28 both times, attribution 22 both times, and a median of 87 both times. That is what you would hope for from checks that are deterministic reads of an HTTP response, and it is the strongest evidence we can offer that these numbers are not an artefact of when we happened to look.
The two that moved are the two that measure the network, which is the honest caveat. Core Web Vitals went from 3 sites failing to 8, and response time from 19 to 21. Both depend on conditions on the day and on Google PageSpeed Insights, so treat those two rows as ranges rather than figures: call it 3 to 8 sites on Core Web Vitals and 19 to 21 on response time. No conclusion on this page rests on either.
We are not in the sample. Our own pages score 97 to 99 on this engine, which is what you would expect from a site built by the people who wrote the checks. That is disclosure, not a result, and it is why visibility100x.com is deliberately excluded from every figure on this page rather than included as a flattering row.
What this study does not show
It does not measure whether these agencies are any good. It measures their own marketing sites. An agency can do excellent work for clients and run a homepage with no date on it, and several in this sample almost certainly do. Read it as a finding about a category's own housekeeping, not as a ranking of suppliers.
It is one page per site. The homepage is the fairest single comparison point across 98 different site structures, and it is still one page. A blog post on the same domain may well carry the date and byline the homepage does not.
The checks are ours. The thresholds for what counts as a pass, a warning and a failure are our engine's, published in the report for every URL it runs on and applied identically to all 98 sites and to our own. Somebody else's thresholds would produce somewhat different counts and, we would expect, the same ordering.
It is a snapshot. Two snapshots, taken a day apart. Sites change, and any of these results can be out of date by the time you read this. The dates are on everything for that reason, and re running it is a documented procedure rather than a favour.
Use this, and cite it like this
The data is published under CC BY 4.0, which means you can republish any of it, including commercially, as long as you credit the source. There is no form, no email wall and no request to link. If you write about this category and need a number, take it.
Visibility100x (2026). Technical audit of 98 generative engine optimization agencies. Published 6 August 2026. https://visibility100x.com/research/geo-agency-audit/
Related: we ran the same engine over the AI visibility tools and found all eleven failing freshness signals on their own sites. What GEO is and what it costs across the market are covered separately.