AI visibility is the easiest thing in this course to measure badly. Answers vary between runs, so a single prompt can tell you almost anything you want to hear. This lesson sets out a method that survives that problem.
Why one prompt proves nothing
Ask the same question three times and you may get three different sets of practices. The causes are structural: generation is probabilistic, location inference varies, live retrieval returns different sources at different moments, and personalisation and conversation history intrude.
So a single favourable answer is not evidence of visibility, and a single unfavourable one is not evidence of failure. Both are draws from a distribution. The only way to see the distribution is repetition.
A fixed prompt set
Write your prompts down once and reuse them unchanged. Changing the wording between measurements destroys comparability. A workable set covers four intents:
- Discovery, unbranded. Who provides a given treatment in your city. This is the question that matters most, because the patient does not know you yet.
- Branded. Ask about your practice by name, to test whether you are described accurately.
- Factual. Address, phone, hours, treatments, insurance. Tests resolution quality.
- Competitive. Ask for alternatives to your practice, to see who the engine groups you with.
Cover each treatment you provide and each city you operate in, and keep the total manageable. Ten to twenty prompts run properly beats a hundred run once.
What to record
For every prompt, on every engine, log:
- Date and engine.
- Whether you were named at all. This is the primary metric.
- Your position in the list, if named.
- Whether the details given were correct.
- Which competitors were named.
- Which sources were cited, where the engine shows them.
- Anything wrong, quoted verbatim, so you can trace it.
Presence rate is the metric to lead with: the share of runs in which you were named. It absorbs variance in a way that position does not.
Method that avoids fooling yourself
- Run each prompt several times and record the presence rate rather than the last answer.
- Use a fresh conversation each time. Context carries over and inflates your results.
- Sign out or use a clean session so personalisation and history do not flatter you.
- Measure on a fixed schedule, monthly is usually right, not when you feel curious.
- Do not change prompts between rounds. If you must add one, add it and keep the originals.
- Record a baseline before making changes, or you will never be able to attribute improvement.
Interpreting the results
- Named in most runs. Working. Monitor and protect it.
- Named occasionally. You are a marginal candidate. Usually improved by better coverage in retrieved sources and stronger prominence.
- Never named. Either you are absent from the sources being retrieved, or resolution is failing. Read the citations to tell which.
- Named with wrong details. The most urgent case, because it actively costs you patients. Trace the source and correct it upstream.
Honesty rules
Three disciplines, particularly if you are reporting this to anyone:
- Never present a single lucky answer as evidence. Report presence rate across runs and say how many runs.
- Never claim a change caused an improvement without a pre-change baseline and a dated change log.
- Say plainly when there is no measurable change. A report that only ever shows gains is not measurement.
These engines are new, variance is high, and the temptation to over-claim is correspondingly strong. A practice that measures honestly knows what is actually working; one that screenshots its best answer knows nothing. Chapter nine applies the same discipline to visibility measurement generally.

