All promptsAnalysis and reporting

Gemini against ChatGPT for SEO work, settled by running the same task in both

A harness you paste unchanged into both assistants. Every claim lands on its own line, tagged sourced, recalled or uncertain, so you judge on your own task.

Works in
ChatGPT, Gemini
You need
One research question you actually need answered · One sentence describing your site and market · Both assistants open, with web search on in each
Written for
gemini vs chatgpt for seo
98A

Scored by our own engine

This page, run through the audit we sell. Measured 4 August 2026.

Score your own page →

Every comparison of these two assistants for SEO arrives at a verdict, and almost none of them show a run. The honest position is that the answer moves with every release and varies by task, so what is worth owning is not a preference but a way of checking. The prompt below is a harness you paste unchanged into both, and it makes each model tag every claim it produces as sourced, recalled or uncertain.

What genuinely differs, and what does not

Both retrieve live web content and both cite what they found. That is the documented, checkable part, and it is a smaller gap than the arguments around it imply.

Google documents grounding with Google Search as connecting the model to real time web content, returning responses with inline citations that link a specific text segment to a source URL. OpenAI describes ChatGPT search in the same shape: answers with links to relevant web sources, with inline citations and a sources panel behind the reply.

Everything past that is behaviour, and behaviour is the thing nobody can hand you a number for. In our own use the two differ mainly in how strictly they hold a defined output format across a long input, and in how readily they decline. Those are tendencies rather than measurements, and they are worth stating as tendencies, because a page that turns them into a score is inventing the score.

Why the tags are the entire test

A retrieved fact and a remembered one look identical on the screen. They arrive in the same register, the same length and the same confidence, and only one of them can be checked.

That is the failure that matters for SEO specifically, because search changes constantly and a model recalling how indexing worked at training time will say so fluently. The tally at the bottom of the output collapses the whole reply into four numbers. Eleven claims and one SOURCED line is a reply built from memory, no matter how well it reads, and that verdict lands on both assistants equally often.

Which one to reach for

Whichever one you have search genuinely switched on in, until your own run says otherwise. That is a duller answer than a ranking and it is the one that survives the next release.

If you are going to standardise on one, run this harness three or four times on questions from your real backlog before you decide, and keep the replies. Then read where Claude fits and where it does not, and take the failure modes seriously with the ways ChatGPT damages SEO. Whichever assistant you settle on, confirm it can read your site at all with the AI crawler checker, because a blocked crawler makes the comparison academic.

The prompt 349 words
You are answering one research question for a search marketer who is going
to paste this same question into a second assistant and compare the two
replies line by line. Write for that comparison.

Answer only the question below. No introduction, no background, no summary,
no closing offer of further help.

Put every claim on its own numbered line, in exactly this shape:

  1. [SOURCED] The claim, in one sentence. (https://example.com/the-page)
  2. [RECALLED] The next claim, in one sentence.

The word in square brackets is the tag, and it is one of exactly three
values:

  SOURCED    you retrieved this from a page during this session and can
             give the URL. Put the URL in brackets at the end of the line.
  RECALLED   this comes from your training data and you have not verified
             it in this session.
  UNCERTAIN  you think it is probably true and would not defend it.

Rules you must follow:

1. Never write SOURCED for a page you did not open in this session. If you
   did not search at all, say so on a line of its own before line 1, and
   then every line below is RECALLED or UNCERTAIN.
2. One claim per line. A line where "and" joins two assertions is two
   lines.
3. Every number, date, percentage, threshold, count and product version is
   a claim. It gets its own line and its own tag.
4. Do not upgrade a RECALLED line to SOURCED by attaching a page that
   merely discusses the subject. The page has to contain the claim.
5. If the question cannot be answered without data you do not have, write
   one line starting NOT ANSWERABLE and name the data, then stop.

After the numbered lines, output these four lines and nothing after them:

  TOTAL CLAIMS: n
  SOURCED: n
  RECALLED: n
  UNCERTAIN: n

Do not comment on the tally. Do not explain it. Do not apologise for it.

The question: [THE ONE QUESTION YOU NEED ANSWERED]
My situation, so you do not answer generically: [ONE SENTENCE ON YOUR SITE AND YOUR MARKET]
Search in this session: [WRITE "on" OR "off"]

What to change

Everything in square brackets is yours to replace. Nothing else needs editing.

[THE ONE QUESTION YOU NEED ANSWERED]
One real question from your own work, not a test question. "Does a canonical pointing at a paginated series still consolidate signals" produces a usable comparison. "What is SEO" produces two essays that agree. Pick something where being wrong would cost you a week.
[ONE SENTENCE ON YOUR SITE AND YOUR MARKET]
What you publish and who reads it, in one line. "We run a 4,000 page recipe site in the UK" changes the answer to almost any indexing question, and without it both assistants answer for a generic ten page brochure site that nobody actually operates.
[WRITE "on" OR "off"]
Whether you have web search enabled in that session. This is the single input that decides what the tally means. Both models will answer either way, and a full page of RECALLED lines with search off is a completely different artefact from a full page of RECALLED lines with search on.

How to run it

  1. 01
    Pick a question where the answer matters

    Use something from your live backlog rather than a benchmark question. The models diverge on the specific and the recent, and they converge on the general, so a broad question produces two near identical replies and tells you nothing about either one.

  2. 02
    Turn web search on in both, and confirm it

    Enable search explicitly rather than trusting the default, then write "on" in the last variable. Search availability differs by plan, by workspace setting and by whether the interface decided the question needed it, so confirm rather than assume.

  3. 03
    Paste the prompt into both, unchanged

    Same wording, same variables, same session start. Editing the prompt between runs is the mistake that makes the whole exercise decorative. Open a fresh conversation in each so neither reply is shaped by something earlier in the thread.

  4. 04
    Read the tally before you read the answer

    The four count lines at the bottom are the finding. A reply with eleven claims and one SOURCED line is a reply assembled from memory, however well written it reads, and that is true of both assistants equally.

  5. 05
    Open two SOURCED links from each and check them

    Click through and confirm the linked page actually contains the claim on that line, which is what rule 4 exists to prevent. This is the step that catches a model attaching a plausible URL to a sentence the page never says.

  6. 06
    Keep the reply with fewer claims, not the longer one

    The shorter reply is usually the more honest one, because most of the extra length in the other is RECALLED filler. Then verify anything you plan to act on against a primary source rather than against the assistant that agreed with you.

Questions people ask

Is Gemini or ChatGPT better for SEO?

Neither has a settled answer, and any page giving you one without showing a run is giving you a preference. Both retrieve live web content and both cite sources when they do. The differences past that are behavioural, they move with every release, and they vary by task, which is why the useful thing to own is a way of testing rather than a verdict.

Do both of them actually search the web?

Yes, and both document it. Google describes grounding with Google Search as connecting the model to real time web content and returning citations that link a text segment to a source URL. OpenAI describes ChatGPT search as answering with links to relevant web sources, with inline citations you can click through to. What neither guarantees is that a given reply used it.

Why does the prompt make the model tag its own claims?

Because the visible difference between a retrieved fact and a remembered one is nothing. Both arrive in the same confident sentence. Forcing a tag onto every line separates them mechanically, and the tally at the bottom turns the whole reply into one number you can read at a glance.

Will the model tag honestly?

Mostly, and not perfectly, which is why step five exists. Rule 4 blocks the common failure, which is attaching a topically related page to a claim that page does not make. Open two of the SOURCED links yourself. If both hold up, the rest of the tally is probably fair.

Can I run this in Claude too?

Yes. It is plain instructions with no vendor specific syntax, so it runs anywhere. The tagging scheme is the point rather than the model, and a three way comparison is more useful than a two way one if you have all three open.

A new prompt, most days One working prompt for a real SEO or AI visibility job, what to change in it, and a worked example. No sequences, no offers dressed as newsletters.

Unsubscribe in one click. We never pass your address on.

Run your first audit
in about a minute

Free account, no card. Paste your URL and get a real, scored report of your AI and search visibility.

Measuring rankings in