Find the content gap between you and a competitor from two sitemaps
Paste two sitemap URL lists and get the subjects your competitor covers and you do not, read from the paths alone, with no guessing at pages nobody pasted.
- Works in
- ChatGPT, Claude, Gemini
- You need
- Your sitemap.xml URL list · Your competitor sitemap.xml URL list · One sentence on what you sell
- Written for
- chatgpt prompt for competitor analysis
Scored by our own engine
This page, run through the audit we sell. Measured 4 August 2026.

A sitemap is a competitor telling you, in public and in machine readable form, every page they thought was worth publishing. This prompt reads two of those lists side by side and returns the subjects they cover and you do not, plus the ones you cover and they do not, which is the half of the exercise everybody forgets to ask for.
What a URL can and cannot tell you
A path is evidence that a page exists on a subject. It is evidence of nothing else.
That single limit is what most competitor gap analysis gets wrong. A slug does not carry word count, publish date, whether the page is indexed, whether anybody reads it, or whether it was written in a burst of enthusiasm two years ago and never touched again. Rule 1 states the boundary and rule 4 enforces it by sending ambiguous slugs to an unreadable list instead of letting the model invent a subject for /p/1183.
The reason to accept that narrow input is that everything inside it is real. You pasted the URLs. The model is grouping evidence rather than recalling an impression of a company it has never seen.
Match the scope or the output is noise
Compare like against like: their blog against your blog, not their whole site against your articles.
The commonest way this run goes wrong is a mismatched scope. Paste your 40 blog posts against their 900 URL sitemap and the gap list fills with product pages, location pages and support documentation, none of which is a content gap. It reads as a devastating competitive deficit and it is a filtering mistake.
Cut both lists to the same section before you run it. If you want the full picture, run it twice with different scopes rather than once with both mixed.
Read your own list first
The strongest finding is usually in “I cover, they do not”, because that is where you already have ground that nobody is contesting. Strengthening three pages there is cheaper than writing fifteen new ones, and it starts from something that already exists.
Before you trust either list, confirm your own sitemap is actually complete with the free sitemap checker, then take the wider view with the four pass competitor analysis or the rest of the analysis prompts.
You are a content strategist. I am giving you two lists of URLs, both taken
from sitemap.xml files: mine and a competitor's. Work out what subjects each
site covers and where the gaps are between them.
You are working from URL paths and nothing else. You have not read a single
one of these pages. You cannot see traffic, rankings, backlinks, publish
dates, page quality, word count or whether any of these pages performs at
all. Do not estimate any of those and do not describe what a page says.
Method:
1. Read a subject from each path. Treat the slug as evidence of what the page
is about and nothing more. "/guides/rota-templates" is evidence that they
have a page about rota templates. It is not evidence that the page is
good, current, indexed or ranking.
2. Group both lists into subject areas. Use the same subject names for both
sites so the two sets can be compared directly.
3. Ignore paths that are not content: pagination, tag and author archives,
search results, legal pages, login and account routes, and anything that
looks machine generated. List what you ignored and why, in one line.
4. Where a slug is genuinely ambiguous, put it in an unreadable list rather
than guessing the subject. A guessed subject is worse than a short list.
5. Do not rank the gaps by opportunity, volume, difficulty or traffic. You
have no data for any of those. You may order them by how close each one
sits to what I said I sell, and you must say that this is what the order
means.
6. If the two sites cover almost the same subjects, say so plainly. A small
gap list is a real result and I would rather have it than a padded one.
Output in this order, with no tables:
THEY COVER, I DO NOT. Each entry: the subject, the competitor paths that show
it, and how many pages they have on it. Group anything they have three or
more pages on into a single line marked "cluster", because a cluster is a
different signal from a single page.
I COVER, THEY DO NOT. Same shape. This is the list people forget to ask for
and it is where the position I already hold is.
BOTH COVER. Subject names only, one per line, no commentary. This exists so I
can see the size of the shared ground.
UNREADABLE. Paths where the subject could not be determined.
WHAT THIS METHOD CANNOT TELL ME. A short list. Include the fact that a
subject appearing in the gap list may be one they tried and abandoned, and
name the one check I would run to find out.
What I sell, and to whom: [ONE SENTENCE ON WHAT YOU SELL AND TO WHOM]
My URLs: [PASTE YOUR SITEMAP URL LIST]
Their URLs: [PASTE THE COMPETITOR SITEMAP URL LIST]What to change
Everything in square brackets is yours to replace. Nothing else needs editing.
[ONE SENTENCE ON WHAT YOU SELL AND TO WHOM]- This is what the gap list gets ordered against. "We sell rota software to independent pub groups" makes a competitor cluster about hotel housekeeping a distant gap rather than an urgent one. Leave it out and every gap looks equally worth filling, which is how a content plan ends up with forty pages on it.
[PASTE YOUR SITEMAP URL LIST]- The URLs from your own sitemap.xml, one per line, with the XML tags stripped. Open the file in a browser, copy the visible list, and paste. If your sitemap is an index pointing at other sitemaps, open the content one rather than the index.
[PASTE THE COMPETITOR SITEMAP URL LIST]- The same from their domain, almost always at /sitemap.xml or named in their robots.txt. This is public information published for crawlers. If they have more than a thousand URLs, paste one section of the site rather than everything, because most models start dropping rows silently past that point.
How to run it
- 01Open both sitemaps and strip the XML
Visit yourdomain.com/sitemap.xml and theirs. If either returns an index file, follow it to the sitemap that lists real pages. Copy the URLs into a plain list, one per line, and drop the lastmod and priority values, which give the model extra fields to reason badly about.
- 02Find the sitemap when it is not at the usual path
If /sitemap.xml returns nothing, open their /robots.txt and look for a Sitemap line, which is where a non standard path is declared. Failing that, try /sitemap_index.xml or /sitemap-index.xml. A site with no sitemap at all is a much smaller list to build by hand than it sounds.
- 03Cut both lists to the same scope
If you are comparing blog coverage, use only blog paths on both sides. Comparing your entire site against their blog produces a gap list made mostly of their product pages, which is not a content gap. Matching the scope is the single step that decides whether the output is useful.
- 04Run it and read the second list first
Start with "I cover, they do not". That list is the ground you already hold and it usually contains the pages worth strengthening rather than the pages worth writing. The gap list is more exciting and more expensive.
- 05Check three gaps by hand before planning anything
Take three subjects from the gap list and open the competitor pages. A gap is only worth filling if the page they have is genuinely serving that subject, and some of them will turn out to be thin pages from a plan they abandoned. The model cannot see this and rule 5 stops it pretending otherwise.
- 06Confirm your own sitemap is complete before you trust the comparison
A gap can be an artefact of your own sitemap missing pages you have published. Run the free sitemap checker on your domain first. A stale or partial sitemap makes you look thinner than you are and sends the whole plan in the wrong direction.
Questions people ask
What is a good ChatGPT prompt for competitor analysis?
One with a defined input and a narrow question. Two sitemap URL lists and the single question of which subjects each site covers is a job a model does well, because the answer is entirely contained in what you pasted. Broad prompts asking for a full competitor analysis have no such input, so the model fills the gap from its training data.
Can ChatGPT read my competitor sitemap if I give it the link?
Sometimes, and it is not worth relying on. Some assistants cannot fetch a URL at all, some return a cached copy, and a large XML file often arrives summarised rather than read in full. Opening the sitemap yourself and pasting the URL list takes a minute and guarantees the model is working from the whole thing.
Is a sitemap content gap the same as a keyword gap?
No. A sitemap gap shows which subjects have a page and which do not, which is a coverage question. A keyword gap shows which queries a competitor ranks for that you do not, which needs rank data from a tool that measures it. The sitemap version is free, immediate and blind to performance.
Why does the prompt refuse to prioritise the gaps?
Because prioritising means comparing opportunity, and opportunity means volume and difficulty, and the model has neither. Asked to rank anyway, it produces an order that looks considered and is arbitrary. Ordering by closeness to what you sell is a judgement it can actually make from the inputs.
What if the competitor has no sitemap?
Check robots.txt for a declared path first, since plenty of sites keep theirs somewhere unusual. If there genuinely is not one, take the URLs from their main navigation and blog index instead and say in the prompt that the list is partial, so the missing subjects are not read as gaps.
Unsubscribe in one click. We never pass your address on.