Measuring your position on Google is (relatively speaking) easy. The search engine is deterministic: the same keyword gives the same result for the same user, every time. The results have explicit rankings. There are APIs that report your position numerically. None of this applies to AI.
Ask ChatGPT the same question twice, five minutes apart, and you’ll get two different answers. The variation isn’t a bug — it’s built into how language models work (through temperature and top-p sampling). A single measurement says little about your actual visibility.
AI models sometimes list options in order, sometimes as flowing text, sometimes implicitly within a sentence. There’s no “position 1” to measure — only patterns across many answers.
There’s no Google Search Console for ChatGPT. No official API that says “your visibility in ChatGPT is X”. The data has to be produced through structured measurement, not pulled from a source.
ChatGPT-4o isn’t the same model today as it was three months ago. When OpenAI updates — and they do so often — your results can change dramatically overnight. A one-off measurement is worthless; only continuous tracking captures reality.
Measuring something no one has measured before requires a disciplined approach. MrSearch is built on three principles we won’t compromise on – because they’re what guarantees that the numbers you get can actually be trusted.
Because AI answers vary, we measure in statistically valid volumes (10 prompts × daily runs) and calculate standard deviation explicitly. We don’t just show a number — we show how reliable that number is.
Every KPI we present must either point directly to an action, or serve as the basis for one. Data points that can’t be translated into “what should I do on Monday morning?” have no place in the product.
When you add a brand, MrSearch uses your URL, the description of your category and your target audience to generate 10 prompts that mirror real search patterns. You can adjust, add or remove them manually. The 10 prompts are the entire measurement instrument — if we choose the wrong prompts, we measure the wrong things.
Each prompt is sent to ChatGPT, Perplexity, Claude, Copilot, Gemini and Google AI Overviews — depending on your plan. The questions are asked neutrally, without personalization, so that the results are comparable between runs and between customers.
Each answer is analyzed by a secondary language model that extracts: which brands are mentioned, in what order, how often, with what sentiment, and which attributes are linked to each one. The results are stored as structured data, not free text.
The extracted data points are fed into ten assessment criteria (more detail in the next section). Each criterion is weighted according to a methodology that is publicly documented. The result is AI-Score: a single number between 0 and 100.
Every data point from every run is stored with a timestamp, model ID and prompt ID. This means you can compare not just today’s AI-Score with last week’s, but exactly which criteria changed — and in which model.
The data is compared against predefined gap patterns (visibility gaps, attribute gaps, position gaps, model gaps). When a pattern matches, an area for improvement is generated with a concrete recommendation and a priority based on the calculated effect on AI-Score.
AI-Score, from 0 to 100, is an average of ten weighted assessment criteria. Here is each criterion, its weight and its rationale — so that you can judge the method for yourself.
Is the brand mentioned at all in the answer? Binary yes/no. The foundation for everything else.
Is the brand the only option or one of several? “Use X” vs “You can choose between X, Y and Z”.
Is the context around the mention positive, neutral or negative? Extracted via secondary API analysis.
How confident is the AI model when it mentions your brand? High probability = strong grounding in the training data.
How early in the answer does the brand appear? Measured in tokens from the start of the answer.
How much text is devoted to describing the brand? Counts tokens directly linked to you.
How many competitors are mentioned in the same answer? Fewer = a stronger standalone position.
How many competitors are mentioned in the same answer? Fewer = a stronger standalone position.
Is the user prompted to visit or use the brand? “Visit X.com” is stronger than a passive name-drop.
Is a URL or source citation included? Relevant mainly for models with web search.
40% of AI-Score is determined by whether you're mentioned and as a primary recommendation. The remaining 60% is about how well you're mentioned — given that you're mentioned.
Two brands can both be mentioned by AI yet have completely different strength in the underlying representation. You won’t notice it in the generated answer. You’ll only notice it if you read the logprobs.
A high logprob for your brand = AI doesn’t have to think in order to recommend you. You’re part of the model’s foundational understanding of the category. It’s the very strongest form of AI visibility there is.
A low logprob = AI mentioned you, but almost by chance. It could just as easily have chosen someone else. Your position is fragile.
The difference between these two is invisible in the ordinary text. It’s visible in the logprobs. And it’s one of the most important signals MrSearch tracks.
The difference in reliability is invisible in the average – visible in the distribution.
avg = #4 · σ = 0.8 (low)
Both have an average position of #4. Only one of them is reliably visible.
Direct API calls to official endpoints (OpenAI, Anthropic, Google, Perplexity). We don’t scrape interfaces. That means our data is the data the model itself produces, not a user-specific, personalized version.
Projects: once per day, all KPIs, all enabled models. One-off measurements (Analyze): in real time. Threshold alerts: triggered immediately when defined limits are exceeded.
The raw responses are stored with a timestamp, model version and prompt ID. This means we can compare historical results even after the models have later been updated — and show you exactly when a shift in visibility coincided with a model update.
Prompts where you’re consistently not mentioned. Indicates a lack of content or authority within the topic area.
Create content that addresses the specific questions. Often product pages, FAQs or comparisons.
Desired brand attributes that the AI models don’t associate with you. Indicates that the right message isn’t present in the content the models are trained on.
Prompts where you’re mentioned but land low in the ranking or as a passive mention. Indicates that competitors have stronger grounding.
Thought leadership, expert citations, third-party mentions. Build authority, not just presence.
AI models where you’re markedly less well represented than in others. Indicates that your content is invisible to that particular model’s data sources.
(and why it matters)
Tracks the performance of your own LLM app (latency, token usage, hallucinations). Useful if you’ve built a chatbot. But it measures your product — not your visibility in external AI services. Completely different problems.
Asking ChatGPT questions yourself is valuable – but subjective, random and unscalable. Your own prompts are also colored by the fact that you already know the answer you’re hoping for. Good for exploration, not for measurement.
Not an SEO tool that “also covers AI”. SEO tools that have started adding AI measurement often do it as an afterthought — low update frequency, shallow KPIs, generic prompts. MrSearch is built from the ground up for AI visibility. It’s our entire product.
The variation is inherent in how language models work. We address it through (1) ten prompts per brand instead of one, (2) daily runs over time, (3) explicit calculation of standard deviation in the results. A number you see in MrSearch is always an average with documented dispersion, not a single snapshot.
Official APIs. OpenAI for ChatGPT, Anthropic for Claude, etc. That means the answers are the model’s neutral output — not a personalized version influenced by a user’s history. It’s a trade-off: the API answers may differ slightly from what an individual user sees in the web interface, but they’re comparable over time and between customers.
We log the model version on every run. When an update is released, it’s marked on your trend lines so you can see exactly when a shift occurred. We make no retroactive adjustments to historical data — it is what it was, measured against the model that existed then.
Prompts can be written in any language. For Swedish brands in Swedish markets we recommend Swedish prompts; for international visibility, prompts are run in English in parallel. You can have both.
We use a secondary large language model (currently Claude and GPT-4o, depending on the task) with structured extraction prompts to categorize brand mentions, sentiment and attributes. The model isn’t specifically trained — it’s used with carefully constructed instructions and is continuously validated against manual annotation.
Yes. CSV export is available today; a public API is on the way.
Yes. We measure publicly available AI answers; no processing of personal data. Your prompts and brand data are stored within the EU and are never used to train external models.
This page is a summary of a product that is considerably more detailed than what fits in text. The easiest way to judge whether MrSearch is right for you is to see it live – on your own category, with your own competitors, in a 20-minute demo.
Cookies and data protection
We use necessary cookies for site functionality and log IP addresses on search requests for security purposes (legitimate interest, GDPR Art. 6(1)(f)). For statistics and marketing we need your consent. Read more in our privacy policy.
Necessary Always on
Session, security log, rate-limit. Always on — required for the site to work.
Statistics
Google Analytics — helps us understand how the site is used.
Marketing
Advertising tracking for targeted campaigns.