GEO & AEO
Written on 19/10/2026
Updated on 19/10/2026
3min

AI visibility audit: 12 copy-paste prompts for Claude, ChatGPT and Perplexity

Thibaut Legrand
Thibaut Legrand
Co-founder - Vydera
AI visibility audit prompts Vydera
Table of contents

An AI visibility audit run on your own data

Vydera logs it, reads the series and fixes it.

Talk to an expert

Key takeaways

  • 12 audit prompts, six axes, two per axis, copied in one click. None of them was run for this article: no model API key, no screenshot, no engine verbatim
  • Over 12 months, 97 queries containing « prompt » reached vydera.com. 43 look for a tool feature: 905 impressions, zero clicks. The other 54 look for a text to reuse: 1,027 impressions, 4.67% CTR
  • 87% of that demand names a model, and Claude carries 875 impressions out of 889. A prompt article naming no engine targets the remaining 13%
  • The reading « naming a model doubles CTR » was tested and rejected. Stratified by position band the gap disappears: it is a composition effect in the traffic, not a wording effect
  • Across 36 sector robots.txt files, zero block a direct-search bot. When an engine ignores your brand, robots.txt is almost never the explanation

The 12 prompts sit further down, in a library with a copy button on each one. Take them, replace the bracketed variables, run them. The rest of this article explains what they measure, how to keep the series alive over time, and why the title names engines.

One thing first, because it changes how you should read everything else: none of these 12 prompts was run for this article. No model API key, no driven browser. So there is no screenshot here, no verbatim from ChatGPT, Claude, Perplexity, Gemini or Copilot, and no measured presence rate. The "good" and "bad" markers attached to each prompt are proposed reading conventions, never observed results.

What is measured, on the other hand, really is: the demand behind the word "prompt" on this site across 12 months of Search Console, and AI crawler access across 40 sector domains. And it is that measurement that changed the article's title halfway through.

Two intents behind the word "prompt", and one never clicks

Over the last 12 months, 97 queries containing the string "prompt" brought impressions to vydera.com: 1,932 impressions, 48 clicks. Averaged out, that gives a 2.48% CTR describing nothing at all, because those 97 queries cover two demands with no relation to each other.

43 of them look for a tool feature: prompt tracking, prompt management, prompt suggestions, platform comparisons. They carry 905 impressions and zero clicks over 12 months. Not a low CTR: a zero on a real denominator. The other 54 look for a text to reuse: 1,027 impressions, 48 clicks, 4.67%.

If you track the word "prompt" in your Search Console, split those two buckets before looking at anything else. The first one calls for a product page or a glossary entry, the second for an article like this one. Merging them halves the CTR and manufactures a figure that describes neither.

The appealing reading, tested and rejected

Inside the "prompt to copy" bucket, the first calculation flatters: queries naming a model show a 5.06% CTR, queries naming none show 2.17%. A ratio of 2.3. The conclusion writes itself, and it would make a good post: name a model in your title and you double your CTR.

It does not hold. Stratified by position band, the difference collapses: 14.02% against 8.33% in positions 1 to 5, 2.89% against 1.92% in 5 to 8, 3.00% against 2.56% in 8 to 11. The group naming no model does not convert worse at equal rank, it simply ranks worse: average position 9.92 against 8.29. It is a composition effect in the traffic, not a wording effect.

And the check stops there, because the cells weigh nothing: 12, 52, 39 and 35 impressions for the group naming no model, band by band. None of them supports a rate comparison. You do not replace a false conclusion with an opposite one that is just as fragile.

What holds: the demand is single-model

What the measurement really establishes is a composition fact, and it is sharp. 889 of the 1,027 impressions on the "prompt to copy" intent contain a model name, that is 87%. And among those 889, Claude carries 875. ChatGPT, Gemini and Perplexity weigh 17 impressions each, all at zero clicks and at average position 70.

Direct consequence for this article. The original brief planned a title with no engine name: "12 prompts to audit your AI visibility". That title does not replicate the format that works, it removes the very word capturing 87% of the demand. Hence the current title, naming Claude first because that is what the site actually receives, and the others next because the protocol runs on five engines.

Second useful cross-check: the qualifier "audit". 16 queries contain it and carry 30 of the 48 clicks, on 437 impressions, against 18 clicks on 590 impressions for the other 38. Careful with the reading: those 16 queries also rank far better, position 5.87 against 10.47. Word and rank are confounded here, and there is no separating them. What can be said, and no more: on this site the productive query is "audit + prompt", not "prompt" on its own.

Before running a prompt: robots.txt is almost never the explanation

When an engine answers "I do not know this brand", the industry reflex is to check robots.txt. On this panel, the reflex is wrong almost every time.

40 sector domains were logged on 27 August 2026, 36 of them with a readable robots.txt. Result: zero block a direct-search bot, meaning OAI-SearchBot, Claude-SearchBot or PerplexityBot. Zero block a user agent either. Five sites block something, and that something is a training crawler or a generic one, CCBot first among them. But blocking CCBot or GPTBot does not remove you from ChatGPT answers: it removes you from a training corpus, which is a different mechanism entirely. The role-by-role detail sits in our AI crawler audit.

The scope of that finding stops at the panel: it says what is common in the sector, not what is true on your site. Check your own file once, it takes thirty seconds. Then move on to the real question, which is editorial.

The 12 prompts, six axes, two per axis

Six axes, because an audit that only asks "does the AI cite me" misses five problems out of six.

  • Brand presence (P01, P02): does the engine name you when nothing tells it you exist?
  • Competitive ranking (P03, P04): where does it rank you, and on which criterion do you lose?
  • Cited sources (P05, P06): which pages does it actually open, and which ones are yours?
  • Factual accuracy (P07, P08): is what it says about your company true, and where does it come from?
  • Topic coverage (P09, P10): across the sub-questions of a topic, how many pages do you have that answer?
  • Brand hallucination (P11, P12): does it invent offerings, prices or an identity you do not have?

What these 12 do that our previous 10 did not

We had published 10 SEO prompts for Claude. It is by far the most read page on the site: the FR and EN versions add up to 899 clicks over 12 months, that is 65.4% of every click measured at page level, and 65.6% against the property total. One instructive detail along the way: the EN version carries 500 clicks and the FR one 399, but the FR page converts twice as well what it receives, 8.9% CTR against 4.3%. More impressions does not mean a better page.

Writing 12 more prompts without saying what they add would be a rerun. So the text of the 10 was read line by line on 28 August 2026, and here are the four gaps counted on it.

  • 4 of the 12 are unaided: the brand name appears nowhere in the instruction. None of the 10 published prompts tests brand presence unaided. The two that look at the brand write its name into the prompt, which turns the test into a confirmation.
  • P02 imposes one turn without web search, then the same turn with it. The 5 web prompts among the published 10 all ask for search from the start, and 4 of them literally open on "Turn on web search". So they always measure retrieval, never what the model knows from memory. The two are repaired differently: memory through durable third-party sources, retrieval through your own pages.
  • All 12 fix their output columns. Two of the published 10 ask for a table, none fixes the columns or requires the full URL. Without columns, two measurements do not compare, and a series does not get built.
  • P08 and P11 carry a control item: a deliberately wrong date, an offering that does not exist. None of the 10 tests what the engine agrees to invent under a question that presupposes. That is the difference between checking what it says and checking what it fabricates.

In total, 6 of the 12 extend an already published prompt and 6 have no equivalent: the substitution test, tracing a claim back to its source, the constrained company record, the contradiction test, the nonexistent product trap and entity disambiguation.

The shape of an audit prompt: long, interrogative, brandless

An audit prompt should not look like a keyword. On the French Search Console of vydera.com, 40 action-shaped queries were served over 12 months, 36 of them 6 words or longer. One example, logged as is: "quels outils geo proposent un accompagnement expert en plus du tracking automatisé ?". That is not what a human types into a search box. It is query fan-out, sub-questions manufactured by an engine to build its answer.

Two guardrails on that corpus. It is single-topic: 11 of the 12 examples are about GEO tooling, the fan-out of one single article on the site. The shape travels, the subject does not, and nothing in there amounts to a general law about query length. And its 244 impressions at zero clicks do not measure the uselessness of those queries: they are strings manufactured by an engine, and the click may simply not exist at that point of the journey.

The logging grid, and its real cost in rows

A prompt run once is an anecdote. What makes an audit is the grid.

Five engines: ChatGPT, Perplexity, Google AI Mode, Claude and Copilot. Three runs per prompt, in a fresh conversation each time. Seven prompts on a monthly cadence, five on a quarterly one. That comes to 105 measurements per monthly campaign, 75 per quarterly campaign, 1,560 rows over a year. The figure is there so nobody commits blind: this is a real workload, not a checkbox at quarter end.

Columns common to every row: date, engine, mode used, prompt id, run number, the exact variable values, and the verbatim sentence carrying the signal. Then the columns specific to each prompt, listed in the library above.

Three rules avoid the most expensive mistakes.

  • A fresh conversation for every run. A thread that keeps context turns the second run into a confirmation of the first, and you end up measuring your own history.
  • Never average a presence rate across engines. The indexes differ, and the average describes no real engine. On how to build that kind of rate, see AI citation rate.
  • Three campaigns before reading a slope. Two points do not make a trend, they make a line.

Reading order matters more than people think. P12 runs first, even though it is numbered last: if your brand is merged with a namesake, the other eleven measurements are mixing two companies and nothing flags it. Then P02, which decides which lever to work. Then P01, P03 and P04 for competitive position, P05 and P06 for sources, which almost always explain the three before them, P07, P08 and P11 for facts and fabrication, and finally P09 and P10 for the editorial plan that follows. On that ground, brand hallucination deserves its own tracking.

Running the demand measurement on your own property

The measured part of this article replays with no paid tool, in four moves.

  1. Export 12 months of queries, with full pagination. A trap verified on our own files: an export capped at 1,000 rows cuts through the middle of the zero-click block, in alphabetical order. On vydera.com it was missing 579 queries and 8,941 impressions, including every accented French query.
  2. Filter on the word, then sort by hand into two buckets. Tool feature on one side, text to reuse on the other. A hundred queries take twenty minutes, and no automatic classifier does better on a corpus that size.
  3. Stratify by position band before comparing two CTRs. It is the only move that separates a wording effect from a rank effect, and it is the one that brought down our nicest hypothesis.
  4. Refuse any cell under a hundred impressions. A rate computed on 12 impressions is not a rate, it is an accident.

What this protocol will not tell you

Better to state it now, before the question lands in a meeting.

  • Three runs do not give a confidence interval. They say whether the answer is stable or not. The run-to-run variance of a single prompt was not measured here, so the number of runs needed for a reliable rate remains unknown.
  • The "good" and "bad" markers come from no observed distribution. They set a convention so two measurements can be compared with each other, nothing more.
  • No recovery delay is promised. No before-and-after series was built, so nobody here can tell you how long a fix takes to show up in a measurement.
  • The number of rows is calculated, the duration is not. 1,560 rows a year, yes. How long that takes, no.
  • The demand figures describe one single site. 97 queries and 1,932 impressions on vydera.com is not the French market for SEO prompts.

Where to start this week

  1. Run P12 on one engine. Ten minutes. If a namesake shows up, stop everything and deal with disambiguation first.
  2. Run P01 in a fresh conversation, three times, on the same engine. Note the rank or the absence, and the ten names cited.
  3. Open a sheet with the common columns plus the ones for those two prompts. That is your baseline, and it is only worth what comes after it.
  4. Add P05 and P06 the following month: sources almost always explain ranks.
  5. Compare nothing before the third campaign.
  • Were these 12 prompts tested on ChatGPT or Claude?

    No, and it is stated in black and white. None of them was run for this article: no model API key, no driven browser. So there is no screenshot and no engine verbatim, and the "good" and "bad" markers are reading conventions, not measured results. What is measured here is the demand and the grid.

  • Should you write your brand name into the prompt?

    It depends what you are after. 4 of the 12 prompts are unaided: the brand appears nowhere in the instruction. A prompt naming the brand tests perception, a prompt that does not tests presence. Both are useful, but they answer different questions, and confusing them makes you believe in a visibility that does not exist.

  • How many times should you rerun a prompt?

    The grid sets 3 runs in a fresh conversation, with no history and no project loaded. Three runs say whether the answer is stable, they give no confidence interval. The run-to-run variance of a single prompt was not measured here: the number of runs needed for a reliable rate stays an open question.

  • Does my robots.txt block AI engines?

    Across 36 sector robots.txt files logged on 27 August 2026, zero block OAI-SearchBot, Claude-SearchBot or PerplexityBot. The 5 sites blocking anything target training crawlers or CCBot, which has no effect on citation in answers. Check yours, but do not lean on it to explain an absent brand. That audit describes a panel, not your site.

  • How many measurements does a full campaign represent?

    105 per monthly campaign: 7 prompts, 5 engines, 3 runs. 75 per quarterly campaign, so 1,560 rows over a year. The duration was not measured: the volume is calculated, the working time is not.

  • Which prompt should you start with?

    P12, entity disambiguation, even though it is numbered last. If your brand is confused with a namesake, the other eleven measurements are mixing two companies and nobody notices. Then P02, which separates what the model knows from memory from what it discovers by searching: the two are repaired with opposite levers.


Thibaut Legrand
Thibaut Legrand
Co-founder - Vydera