How we measure something no one has been able to measure before

AI answers are non-deterministic. The models are updated every week. There are no official rankings. This is how MrSearch produces measurable, repeatable and actionable numbers on your AI visibility anyway — every day.
This page is technical. It's written for those of you who want to understand the methodology before relying on the numbers. Jump straight to a section using the menu below, or read all the way through.

Why AI visibility is fundamentally different from SEO

Measuring your position on Google is (relatively speaking) easy. The search engine is deterministic: the same keyword gives the same result for the same user, every time. The results have explicit rankings. There are APIs that report your position numerically. None of this applies to AI.

Non-deterministic answers

Ask ChatGPT the same question twice, five minutes apart, and you’ll get two different answers. The variation isn’t a bug — it’s built into how language models work (through temperature and top-p sampling). A single measurement says little about your actual visibility.

No explicit ranking

AI models sometimes list options in order, sometimes as flowing text, sometimes implicitly within a sentence. There’s no “position 1” to measure — only patterns across many answers.

Dark data

There’s no Google Search Console for ChatGPT. No official API that says “your visibility in ChatGPT is X”. The data has to be produced through structured measurement, not pulled from a source.

The models are updated continuously

ChatGPT-4o isn’t the same model today as it was three months ago. When OpenAI updates — and they do so often — your results can change dramatically overnight. A one-off measurement is worthless; only continuous tracking captures reality.

These are the four challenges MrSearch is built to solve. The next sections explain exactly how.

Three principles that govern everything we measure

Measuring something no one has measured before requires a disciplined approach. MrSearch is built on three principles we won’t compromise on – because they’re what guarantees that the numbers you get can actually be trusted.

1
Rigor

"No number may exist without a transparent method behind it."

Every KPI we show has a documented calculation method, weighted criteria and defined data sources. You should be able to ask us exactly how a number was arrived at — and we should be able to answer.
2
Repeatability

"Same method, same data, same result."

Because AI answers vary, we measure in statistically valid volumes (10 prompts × daily runs) and calculate standard deviation explicitly. We don’t just show a number — we show how reliable that number is.

3
Actionable insights

"Measurement without action is meditation."

Every KPI we present must either point directly to an action, or serve as the basis for one. Data points that can’t be translated into “what should I do on Monday morning?” have no place in the product.

From your URL to actionable insight - in six steps

Everything MrSearch does runs through the same pipeline. Here is every step, honestly described — including what happens in the background that you don’t see in the interface.
1

Generation

Prompts

10 prompts are created from your URL and category
2

Query

Multi-model

The prompts are run against all AI models daily
3

Analysis

Answer analysis

Structured extraction of data from the answers
4

Scoring

Weighting

10 criteria are weighted into AI-Score
5

Storage

Trends

Timestamped data for history and comparison
6

Output

Actions

Prioritized recommendations based on gaps

"What would your customers actually ask AI?"

When you add a brand, MrSearch uses your URL, the description of your category and your target audience to generate 10 prompts that mirror real search patterns. You can adjust, add or remove them manually. The 10 prompts are the entire measurement instrument — if we choose the wrong prompts, we measure the wrong things.

"Same questions, all relevant AI models, every day."

Each prompt is sent to ChatGPT, Perplexity, Claude, Copilot, Gemini and Google AI Overviews — depending on your plan. The questions are asked neutrally, without personalization, so that the results are comparable between runs and between customers.

"What did AI actually say — and who was mentioned?"

Each answer is analyzed by a secondary language model that extracts: which brands are mentioned, in what order, how often, with what sentiment, and which attributes are linked to each one. The results are stored as structured data, not free text.

"Translate the raw data into a comparable number."

The extracted data points are fed into ten assessment criteria (more detail in the next section). Each criterion is weighted according to a methodology that is publicly documented. The result is AI-Score: a single number between 0 and 100.

"Not just today's number — the history behind it."

Every data point from every run is stored with a timestamp, model ID and prompt ID. This means you can compare not just today’s AI-Score with last week’s, but exactly which criteria changed — and in which model.

"What should you do about it?"

The data is compared against predefined gap patterns (visibility gaps, attribute gaps, position gaps, model gaps). When a pattern matches, an area for improvement is generated with a concrete recommendation and a priority based on the calculated effect on AI-Score.

AI-Score is a single number. Behind it stand ten.

AI-Score, from 0 to 100, is an average of ten weighted assessment criteria. Here is each criterion, its weight and its rationale — so that you can judge the method for yourself.

1

Mentions

Weight

Is the brand mentioned at all in the answer? Binary yes/no. The foundation for everything else.

2

Primary recommendation

Weight

Is the brand the only option or one of several? “Use X” vs “You can choose between X, Y and Z”.

3

Sentiment

Weight

Is the context around the mention positive, neutral or negative? Extracted via secondary API analysis.

4

Logprobs / Model confidence

Weight

How confident is the AI model when it mentions your brand? High probability = strong grounding in the training data.

5

Text position

Weight

How early in the answer does the brand appear? Measured in tokens from the start of the answer.

6

Attribute density

Weight

How much text is devoted to describing the brand? Counts tokens directly linked to you.

7

Exclusivity

Weight

How many competitors are mentioned in the same answer? Fewer = a stronger standalone position.

8

Mention frequency

Weight

How many competitors are mentioned in the same answer? Fewer = a stronger standalone position.

9

Call to action

Weight

Is the user prompted to visit or use the brand? “Visit X.com” is stronger than a passive name-drop.

10

Source citation

Weight

Is a URL or source citation included? Relevant mainly for models with web search.

Total 100%
AI-Score 10 CRITERIA
Weight distribution

40% of AI-Score is determined by whether you're mentioned and as a primary recommendation. The remaining 60% is about how well you're mentioned — given that you're mentioned.

The weighting isn't random

Mentions and Primary recommendation together make up 40% — because if you're not mentioned at all, or only mentioned as one of many options, the rest of the criteria matter only marginally. The remaining 60% is about how well you're mentioned, given that you're mentioned.

Some criteria are deliberately low

Source citation (1%) will carry more weight in the future, when web search becomes standard in all AI models. We weight for the landscape that exists today, not the one we think will exist tomorrow.

We don't just measure what AI says - but how confident it is about it

When a language model generates an answer, it doesn’t just produce the text — it also produces, for each word, the probability that that particular word would be chosen. That probability is called log-probability, or logprobs. It’s one of the most underused signals in all of AI-era measurement.

Why logprobs are interesting

When ChatGPT writes “One of the best solutions is Acme”, the word “Acme” can have:

High logprob (95%) = AI is confident.

Acme is strongly grounded in the model's training data for this type of question. It's not a guess, it's an established association.

Low logprob (12%) = AI is uncertain.

The model chose this word, but there were several other equivalent options. The association is weak.

Two brands can both be mentioned by AI yet have completely different strength in the underlying representation. You won’t notice it in the generated answer. You’ll only notice it if you read the logprobs.

What it means for you

A high logprob for your brand = AI doesn’t have to think in order to recommend you. You’re part of the model’s foundational understanding of the category. It’s the very strongest form of AI visibility there is.

A low logprob = AI mentioned you, but almost by chance. It could just as easily have chosen someone else. Your position is fragile.

The difference between these two is invisible in the ordinary text. It’s visible in the logprobs. And it’s one of the most important signals MrSearch tracks.

// Example: Two answers, same question, different confidence
One of the best solutions for fast file management is Acme which is known for stability .
Logprob for "Acme": 0.94 (very high)
// Same question, different brand:
One of the best solutions for fast file management is Beta which offers similar features .
Logprob for "Beta": 0.31 (low)
low
high

How we get statistically valid numbers out of random answers

Language models produce different answers to the same question. That’s a fact there’s no getting around. But it doesn’t mean they can’t be measured — it just means the measurement has to be designed to handle the variation.
10
Prompts / Brand
Ten prompts, not one
A single prompt can give a misleading result because of the model’s randomness. Ten prompts covering different ways of asking about the same category produce an average that dampens the variation. Ten isn’t a magic number — it’s a trade-off between statistical strength and measurement cost.
24h
Update cycle

Daily run

Every prompt is run every day. Trends across days dampen individual outliers. A change that holds for seven days is signal; a single day’s variation is noise.
σ
Dispersion measured

Standard deviation, explicitly

In addition to the average, we calculate the dispersion in your results. A brand that lands at position 2 in 9 of 10 prompts (low σ) is more grounded than one that lands at position 1 in 5 prompts and position 8 in 5 prompts. Both have an average position of around 4–5; only one of them is consistently visible.
Two brands. Same average position. Different consistency.

The difference in reliability is invisible in the average – visible in the distribution.

Brand A — consistent

avg = #4 · σ = 0.8 (low)

Brand B — scattered

avg = #4 · σ = 2.7 (high)

Both have an average position of #4. Only one of them is reliably visible.

Where the data comes from - and how it's stored

Data sources

Direct API calls to official endpoints (OpenAI, Anthropic, Google, Perplexity). We don’t scrape interfaces. That means our data is the data the model itself produces, not a user-specific, personalized version.

Frequency

Projects: once per day, all KPIs, all enabled models. One-off measurements (Analyze): in real time. Threshold alerts: triggered immediately when defined limits are exceeded.

Storage

The raw responses are stored with a timestamp, model version and prompt ID. This means we can compare historical results even after the models have later been updated — and show you exactly when a shift in visibility coincided with a model update.

Data privacy
Your prompts and brand data are your property. We don’t use them to train internal models. The data is stored within the EU.
Your prompts
MrSearch Orchestrator
ChatGPT
Claude
Perplexity
Copilot
Gemini
AIO
Answer analysis (secondary LLM)
Structured data (EU storage)
Your dashboard

How a KPI change becomes a concrete action

Language models produce different answers to the same question. That’s a fact there’s no getting around. But it doesn’t mean they can’t be measured — it just means the measurement has to be designed to handle the variation.

01 - Visibility gap

Prompts where you’re consistently not mentioned. Indicates a lack of content or authority within the topic area.

Action pattern

Create content that addresses the specific questions. Often product pages, FAQs or comparisons.

02 - Attribute gap

Desired brand attributes that the AI models don’t associate with you. Indicates that the right message isn’t present in the content the models are trained on.

Action pattern
Publish content that explicitly links the attribute to the brand — case studies, product pages, comparisons.

03 - Position gap

Prompts where you’re mentioned but land low in the ranking or as a passive mention. Indicates that competitors have stronger grounding.

Action pattern

Thought leadership, expert citations, third-party mentions. Build authority, not just presence.

04 - Model gap

AI models where you’re markedly less well represented than in others. Indicates that your content is invisible to that particular model’s data sources.

Action pattern
Identify the model’s specific sources (Reddit, Wikipedia, industry forums) and build a presence there.
Each gap is weighted by its potential impact on AI-Score. The one with the highest calculated effect is shown first. You don't get 50 data points to sort through — you get 3–5 prioritized actions.

What MrSearch is not

(and why it matters)

Alternative 01

LLM monitoring tools

Tracks the performance of your own LLM app (latency, token usage, hallucinations). Useful if you’ve built a chatbot. But it measures your product — not your visibility in external AI services. Completely different problems.

Alternative 02

Manual prompt testing

Asking ChatGPT questions yourself is valuable – but subjective, random and unscalable. Your own prompts are also colored by the fact that you already know the answer you’re hoping for. Good for exploration, not for measurement.

Alternative 03

Built solely for AI visibility

Not an SEO tool that “also covers AI”. SEO tools that have started adding AI measurement often do it as an afterthought — low update frequency, shallow KPIs, generic prompts. MrSearch is built from the ground up for AI visibility. It’s our entire product.

Questions from CTOs, SEO managers and data engineers

The variation is inherent in how language models work. We address it through (1) ten prompts per brand instead of one, (2) daily runs over time, (3) explicit calculation of standard deviation in the results. A number you see in MrSearch is always an average with documented dispersion, not a single snapshot.

Official APIs. OpenAI for ChatGPT, Anthropic for Claude, etc. That means the answers are the model’s neutral output — not a personalized version influenced by a user’s history. It’s a trade-off: the API answers may differ slightly from what an individual user sees in the web interface, but they’re comparable over time and between customers.

We log the model version on every run. When an update is released, it’s marked on your trend lines so you can see exactly when a shift occurred. We make no retroactive adjustments to historical data — it is what it was, measured against the model that existed then.

Prompts can be written in any language. For Swedish brands in Swedish markets we recommend Swedish prompts; for international visibility, prompts are run in English in parallel. You can have both.

We use a secondary large language model (currently Claude and GPT-4o, depending on the task) with structured extraction prompts to categorize brand mentions, sentiment and attributes. The model isn’t specifically trained — it’s used with carefully constructed instructions and is continuously validated against manual annotation.

Yes. CSV export is available today; a public API is on the way.

Yes. We measure publicly available AI answers; no processing of personal data. Your prompts and brand data are stored within the EU and are never used to train external models.

Want to see the method in action?

This page is a summary of a product that is considerably more detailed than what fits in text. The easiest way to judge whether MrSearch is right for you is to see it live – on your own category, with your own competitors, in a 20-minute demo.