Blog

Why Your AI Visibility Score Moves When Nothing Changed

17 September 2026

You changed nothing on your site, and your AI visibility score went up six points. Or down. Before you credit a new page or blame a competitor, it is worth knowing that the instrument itself moves, in ways that are easy to miss because nothing breaks.

We run weekly scans across ChatGPT, Perplexity, Gemini, Claude and DeepSeek. Three sources of movement have shown up in our own data, and none of them is about your brand.

The model behind the name can change

What we found on 1 September

Every AI API lets you ask for a model by name. We checked what each engine actually served against what we asked for. Three matched. Two did not:

EngineModel we asked forModel that answered
ChatGPTgpt-4ogpt-4o-2024-08-06
DeepSeekdeepseek-chatdeepseek-v4-flash

The ChatGPT one is harmless: the short name points at a dated snapshot of the same model. The DeepSeek one is not. A different model generation had been answering under the old name since some date we could not find in any record, and every answer still looked normal.

Then it happened again, with the same name

On 10 September DeepSeek released a new version and gave it a name with no version number in it. The old name kept working and quietly routed to the new model.

Our scans on 14 September were the first after that change. So that week's DeepSeek comparison was a comparison between two different models, and the "change since last week" for that engine measured the model, not the brand.

What we do about it now

Every scan now records, on our side, which model was requested and which one answered, and alerts us to any difference. For DeepSeek, where the name no longer changes, we also record a fingerprint the API returns for the serving system, so a new model under an old name still shows up. For Claude we deliberately ask for a fixed, dated model, because a loud retirement is better than a silent swap.

The same question does not get the same answer

One setting doesn't fix it

The usual advice is to set "temperature" to 0, which is meant to make a model pick its most likely answer every time. We tested that on one of our own internal checks, with the input identical byte for byte.

The model gave the same answer 12 times in a row. Fifteen minutes later it gave a different answer, 8 times in a row. Within one session it looked perfectly stable. Across sessions it was not. Several newer models no longer accept the setting at all.

Why that matters for a score

If a single check can flip between sessions, then one answer per question is a coin you flipped once. This is why we ask every tracked question in six different phrasings, on five engines: 30 answers per question per week. A brand named in 4 of 30 and then 6 of 30 has moved a little. A brand named in 0 of 1 and then 1 of 1 has moved from "never" to "always" on a single roll.

The web underneath keeps changing

Engines that search are reading live pages

Four of the five engines we track search the web before they answer: ChatGPT, Perplexity, Gemini and Claude. Their answers are built from pages fetched at that moment. A new "best tools" article, a forum thread that starts ranking, or a page that drops out of an index can change who gets named. That is a real change in the world, but it is not always a change you made.

This one is worth watching, not ignoring

Unlike the first two, this movement tells you something about your market. When a competitor starts appearing where they did not before, ask that engine the question yourself and look at the sources it cites. That is often where the next piece of work is.

How to read a weekly score without being fooled

Look at more than one week

A single week's change is weak evidence. Two or three weeks moving the same way is a trend. We show the history for this reason, and we would rather you ignore one odd week than act on it.

Check whether the engines agree

If your score rose because one engine started naming you and the other four did not change, look at that engine first. Did its model change that week? A move that shows up across several engines is much harder to explain away.

A worked example

Say you were named in 9 of 30 answers last week and 13 of 30 this week. Before you credit the new comparison page you shipped, check three things. First, did any engine change its model that week? Providers announce some of these changes and not others. Second, did the gain come from one engine or several? Third, when you ask the question yourself, do the cited sources include the page you changed?

If the four extra answers all came from DeepSeek in the week its model changed, you have learned about DeepSeek. If they came from three engines and two of them cite your new page, you have learned about your page.

What this means for any AI visibility number

Every tool in this category, ours included, is measuring systems that change without notice. The honest version of a score says what it was measured with. When you compare tools or read a report, ask three questions: how many answers per question, which models answered, and what the tool does when those models change.

If you want to see these numbers for your own brand, the no-signup scan gives a first reading in about a minute. A free account tracks the same questions every week, so one odd week sits next to the ones around it.