Use Perplexity to find out who is currently cited for your buying questions
A prompt built for a cited answer engine: it returns the pages currently quoted for your category questions, and the pattern in why those pages and not yours.
- Works in
- ChatGPT, Claude
- You need
- Your category and what you sell · Five questions your buyers ask · Your domain
- Written for
- perplexity prompt
Scored by our own engine
This page, run through the audit we sell. Measured 8 August 2026.

Answer engines that cite their sources publish something ordinary search does not: the actual list of pages that were good enough to use. This prompt treats that list as the data. It asks five buying questions, then reports which domains were used, of what type, and what the used pages had in common.
Ask what got cited, not what the answer is
The instinct is to ask an answer engine your question and read the answer. The answer is the least interesting output.
The source list underneath it is a live sample of which pages a retrieval system judged good enough to build an answer from, for a question your buyers actually ask. That is a direct measurement of something most people infer from rank data, and it is available for free, several times a day, in a tool most marketing teams already have open.
The pattern section, and why it is written that way
Section B insists on properties rather than adjectives, and the constraint is doing real work.
Ask any model what cited pages have in common and the default answer is that they are authoritative, comprehensive and high quality. That is true, unfalsifiable, and impossible to act on. Forcing the answer into things a page either has or does not have produces a different kind of list: publishes dated prices, includes a comparison table, states a method next to a figure, carries a visible last updated date, describes a procedure in numbered steps.
Those are jobs. You can do them this week and check afterwards whether the citation picture changed.
Absence usually has a boring cause
When a domain appears in none of the five runs, the interpretation people reach for is a content problem. Check the mechanical explanation first.
Retrieval agents can be blocked in robots.txt, at the CDN, or by a firewall rule nobody remembers adding, and a blocked site is invisible regardless of how well it is written. We find this often enough that it is the first thing to rule out, and it takes five minutes with the AI crawler checker. The related failure is a page whose content only exists after JavaScript runs, which the render gap checker catches.
Only once both are clear is absence a finding about your content.
What happened when we ran it without retrieval
We tested this prompt on 8 August 2026 against two models with web access switched off, to see what the failure actually looks like. The two behaved completely differently and both results are worth knowing.
Claude refused. It said plainly that it could not browse, could not verify that pages still exist, could not confirm publication dates, and therefore could not complete the task as specified. That is the correct answer.
ChatGPT produced the entire report. Five questions answered, every source classified by type, a reason attached to each, and a tidy citation list with URLs. It looked exactly like a successful run. We opened the links: one of them, an Atlassian page on project management adoption pitfalls, returns a plain 404, because the page does not exist and never did. Several others could not be verified either way.
This is the whole reason step two exists. A run without retrieval does not fail loudly, it fails beautifully, and the output is indistinguishable from a good one until you start opening links. If you cannot confirm your tool retrieved live sources, throw the run away rather than reading it.
What two runs buy you
The single most common misreading of any exercise like this is treating one run as a measurement. These systems are not deterministic, and the same question asked twice returns overlapping but different source sets.
Two runs in separate conversations, with the same five questions, is the minimum that distinguishes an established source from a lucky one. It doubles the effort of a twelve minute exercise, which is a good trade for the difference between data and an anecdote.
For the broader version of this, run across several assistants and tracked over time, use the AI visibility audit prompt. To understand why a source can be cited without ranking at all, how to rank in Google AI Overviews covers the mechanism.
Answer using current web sources and cite every claim.
I want to understand which pages are being used as sources for questions in
my category, and what those pages have in common.
My category: [CATEGORY]
What I sell: [WHAT YOU SELL]
My domain: [YOUR DOMAIN]
For each of the questions below, do this:
1. Answer the question as you normally would, briefly.
2. List every source you drew on, with the publisher and the page title.
3. For each source, classify it: vendor site, independent review or
comparison, community thread, news or trade publication, documentation,
research or study, or video.
4. For each source, say in one clause what made it usable for this question.
Choose from: it contained a specific figure, it named prices, it listed
options in a structured way, it described a procedure step by step, it
carried a recent date, it was the primary source others cite, or it was
the only page addressing this at all.
The questions:
[FIVE QUESTIONS]
After you have done all five, give me:
A. A table of every domain you cited across all five questions, with how
many of the five it appeared in, and its dominant source type.
B. The pattern: what the frequently cited pages have in common, stated as
things a page either has or does not have, not as adjectives. "Publishes
dated prices in a table" is useful. "High quality and authoritative" is
not.
C. Whether my domain appeared for any of the five, and if it did not,
which single question it was closest to being relevant for.
D. The two source types doing the most work in this category, and whether a
vendor page can realistically compete for those questions or not.
Rules:
- Cite everything. If you cannot find a source for a claim, leave the claim
out rather than asserting it.
- Do not include my domain in any list unless a source you actually used
came from it. Do not be generous to me.
- Where a page you cite is undated, say undated rather than estimating when
it was written.
- Do not tell me how to improve my site. I am asking what is there, not
what to do about it. Recommendations come after the picture is accurate.What to change
Everything in square brackets is yours to replace. Nothing else needs editing.
[CATEGORY]- The category as a buyer would name it, not as your marketing does. If your positioning language and the words buyers use have drifted apart, use theirs, because theirs is what gets typed and spoken into an assistant.
[WHAT YOU SELL]- One sentence. It exists so section C can judge whether your domain was even close to relevant for a question, rather than just reporting that it did not appear.
[YOUR DOMAIN]- Your bare domain. The instruction not to be generous to you matters here: without it a model will sometimes list your site as a source because you asked about it, which turns the one measurement you cared about into a false positive.
[FIVE QUESTIONS]- Five questions a buyer asks before choosing, written as they would speak them. The good ones are the awkward ones: what does this actually cost, what goes wrong with it, what do people switch to, is it worth it for a small team. Questions phrased the way your category page is titled will return your category page and teach you nothing.
How to run it
- 01Write the five questions from real buyer language
Take them from sales calls, support tickets or a community your buyers use. Questions invented at a desk tend to be the ones your marketing already answers, which is why the exercise so often returns a reassuring and useless picture.
- 02Run it in an answer engine with live retrieval
This prompt is built for a tool that retrieves and cites in real time. Run in a model without retrieval it will produce sources from memory, some of which will not exist, and the citation table is then worse than no data because it looks like data.
- 03Run it twice, in separate conversations
Answers are not deterministic and the source set moves between runs. Two runs tell you which domains are consistently used and which showed up once. A domain appearing in both is a finding. A domain appearing in one is a coin flip.
- 04Read section B and ignore everything vague in it
The pattern section is where the value is, and only the concrete parts of it count. Dated prices, a comparison table, a named procedure, a figure with a method attached: these are things you can go and add. Anything that reads as a compliment about quality is the model padding and should be discarded.
- 05Check that assistants can reach you at all
If your domain appeared nowhere, confirm the boring explanation before the interesting one. Run the AI crawler check on your site to see whether the retrieval agents are even allowed to fetch it. A robots rule is a far more common cause of total absence than a content problem, and it is a five minute fix.
Questions people ask
What makes a good Perplexity prompt?
One that asks for something the citations can prove, since a cited answer engine has a real source list attached to every answer. Asking it to reason at length wastes what makes it different. Asking it what it used, of what type, and what those pages had in common turns the citation list itself into the data.
Can I use this prompt in ChatGPT or Claude instead?
Only with web search or browsing switched on. Without live retrieval the model will supply a source list from memory, and those citations are frequently plausible and wrong. The prompt depends entirely on the sources being real, so a run without retrieval should be discarded rather than interpreted.
Why run it twice?
Because these systems are not deterministic and the retrieved sources vary between runs for the same question. A single run tells you what happened once. Two runs separate the domains that are genuinely established for a question from the ones that appeared by chance, and that distinction is the whole point of the exercise.
What if my domain does not appear at all?
Check whether you are reachable before concluding anything about your content. Assistants use retrieval agents that can be blocked in robots.txt or by a firewall or CDN rule, often without anybody deciding to, and a blocked site cannot be cited no matter how good it is. If access is fine, then absence is a content and authority finding and the pattern in section B is the brief for fixing it.
Is this the same as rank tracking?
No. Rank tracking measures position for a phrase. This measures which pages get used as sources for a question, which is a different set: a page can be cited without ranking well, and a page can rank first and never be used. Both are worth knowing and they answer different questions.
How often should I rerun it?
Quarterly is enough for most categories, with the same five questions each time so the runs are comparable. The value is in the change: a domain that appears in three of five this quarter and one of five next quarter is telling you something that a single snapshot never could.
Unsubscribe in one click. We never pass your address on.