Methodology

How we measure what AI says about a business.

AI answers vary between models, between days and between ways of asking. A measurement is only useful if it says exactly what was asked, of which model, how it was counted and what is unknown. This is how we do it.

In short

We ask a frozen panel of real buyer questions to ChatGPT, Claude, Gemini, Perplexity and Grok with live web search, keep every answer with its cited sources, and report — per model — how often each business is named, cited, presented favourably and recommended for complex needs, with counts, totals and unknowns.

The AI models

Which models we test, and how.

We test through each provider's official API with live web search switched on, and record the exact model version that answered.

API answers are not identical to the consumer apps. The ChatGPT, Claude, Gemini, Perplexity and Grok apps can add personalization, location and other settings. We report each model as its own series and never present one model as “AI” as a whole.

AI models we test, as of October 2026
AI assistantStandard checkDeep check adds
ChatGPT (OpenAI)GPT-5.6 SolGPT-5.6 Luna, GPT-6 Luna, GPT-6.1 Sol, GPT-6 Astra
Claude (Anthropic)Claude Sonnet 5.5Claude Haiku 4.5, Claude Opus 5.5, Claude Fable 5.1
Gemini (Google)Gemini 3.8 FlashGemini 3.6 Flash, Gemini 3.1 Pro
PerplexitySonar ProSonar, Sonar Reasoning Pro, Sonar Deep Research
Grok (xAI)Grok 4.6Grok 4.7, Grok 4.6 Heavy
Not covered yetGoogle's AI Overviews and AI Mode, Microsoft Copilot and Meta AI. Experimental models are reported separately.

The questions

Real buyer questions, frozen so results compare.

The question panel is built from your market — every product or service you and your competitors offer — not from your marketing copy.

Broad

Discovery questions

“Which businesses near me offer bookkeeping for restaurants?” — Does the assistant know you exist for what you sell?

Specific

One stated need

“Which bookkeepers near me publish their prices?” — Are you named when the buyer asks about one fact?

Complex

Several requirements at once

“Which bookkeepers near me handle payroll for 20 staff and say what is included?” — Are you recommended when the buyer is close to choosing?

Brand

Questions that name you

“How does your firm compare with a named competitor?” — Reported separately, so asking about your own name never inflates discovery results.

Questions are written without your brand name unless they are brand questions. The core panel stays the same from month to month; new questions are tracked as a separate panel.

The metrics

Four outcomes, each with a clear denominator.

Every rate is shown with its count and total. Answers that failed, were cut off or could not be classified are reported as unknown — never silently counted as a “no”.

Mention rate

Is the business named at all? An exact name or website must appear in the answer; a similar-sounding name is not counted.

answers naming the business ÷ complete answers

Own-page citation rate

Is one of the business's own pages a cited source? A mention is not a citation; we record both, and which statement each citation supports.

answers citing an owned page ÷ complete answers

Favourable inclusion

Is the business offered as an option, presented positively or neutrally — not named only as a warning or an example of what not to choose?

answers including it favourably ÷ complete answers

Complex-question recommendation

When a buyer states several requirements, is the business recommended as a fit? Only questions it is eligible for count, based on what its pages state.

qualified recommendations ÷ eligible complex answers

The readiness score

What the readiness score means — and what it doesn't.

The readiness score measures your published evidence: how clearly and completely your pages state the facts buyers need to answer each question in your market.

  • Calculated for the business, each offering and each page
  • Every point traced to a passage, or to the pages we examined when a fact is missing
  • Facts we could not check are excluded and the score is marked provisional
  • No points for formatting tricks or markup that the page does not back up

The score is not a prediction of AI behaviour. It describes what your pages state. We report it next to the AI outcomes we measured, not as their cause.

We are building the evidence to test how well readiness relates to AI outcomes across many businesses, markets and models. Until that evidence passes our own validation threshold, every report says the score is unvalidated — and we will not claim otherwise.

Limits

What any GEO measurement can and cannot prove.

What it shows

  • How specific AI models answered specific questions, on specific dates
  • Which pages and sources those answers relied on
  • How you compare with competitors on the same questions
  • Whether results moved after changes, per model

What it can't show on its own

  • Exactly what every person sees in every AI app
  • That a change caused a result, without a comparison group
  • Revenue impact, unless leads are tracked separately
  • How a model will behave after its next update

Next step

See these numbers for your own business.

An audit applies this method to your market: your questions, your competitors, every AI model we test.