How Zene measures AI visibility
Which engines we ask, how we detect your brand, how the score is calculated, and where the numbers stop being reliable.
AI answers are not a ranked list you can check once. They change between runs, differ by engine and country, and depend on whether the model searched the web. This page documents how Zene turns that into numbers you can track, written from the code that computes them. When the method changes, this page changes with it.
The short version
- We ask each AI engine your buyers' real questions through its official API. Your brand name is never in the prompt.
- Each answer is checked for your brand and its aliases. We record whether you were named, in which position, and whether the tone was positive, neutral or negative.
- Each answer earns up to 1 point: mention × tone weight × position weight. An engine's score is the average over your tracked questions, on a 0-100 scale.
- Your overall score is the average of your engines' daily scores over the last 7 days that have data.
- Mention rates come with a 95% confidence range, because a handful of answers cannot prove much.
- Visibility alerts only fire when two consecutive full scans agree.
Which engines we ask, and how
Zene scores 5 AI engines. Each one is called through its provider's official API with a fixed, cost-efficient model. Whether an engine searches the web before answering matters a lot: web-grounded answers can react to new content within days, while knowledge-only answers reflect what the model learned in training.
| Engine | Called through | Live web search | How the market is applied |
|---|---|---|---|
| ChatGPT | OpenAI API | On for every call. The web search tool is required, not optional. | Country passed as the searcher's approximate location |
| Claude | Anthropic API | On. Anthropic's web search tool is attached to every scan call. | Country passed as the searcher's approximate location |
| Perplexity | Perplexity Sonar API | Always on. Live web search is built into the engine. | Country passed as the searcher's location |
| Gemini | Google Gemini API | Off. Answers come from the model's own knowledge. | Outside the US, a one-line note says which country the person asking is in |
| Grok | xAI API | Off. xAI's Live Search is not enabled, so answers come from model knowledge and carry no citations. | Outside the US, a one-line note says which country the person asking is in |
Why is Gemini's web search off? Google's search grounding cannot be combined with the structured output we use to read the list of brands reliably, so today we keep brand detection accurate and accept knowledge-only answers. These settings are part of our versioned scan configuration and can change; when they do, this page changes too.
The Free plan scans ChatGPT and Gemini. The paid plans add Claude, Perplexity and Grok. On a paid plan, Zene can also capture the answer from ChatGPT's web interface as a separate check; those captures are kept apart from the score.
Questions and markets
You choose the questions Zene tracks. We suggest brand-agnostic ones, the way a buyer would actually ask ("best CRM for a small agency", not "is Acme a good CRM"), because those are the answers where a brand has to earn its place.
Prompts are brand-blind. The engine receives the buyer's question and nothing about you. It answers the way it would for anyone, then lists every brand it named, in order, with a tone label. Naming your brand in the prompt would nudge the model to include you and inflate the number we are trying to measure. Answers are requested in the question's language and kept to about 250 words, so the list of brands is not cut off.
Every question is asked in a market (a country). By default that is your brand's home country; brands without one are scanned as the United States. The Free plan covers 1 market, Pro up to 1 and Agency up to 10. How the country reaches each engine is shown in the table above. Location is country-level, never city-level.
Large question sets rotate. When a brand tracks more questions than one scan covers, each scan takes the questions that were scanned least recently first, so every question is refreshed in turn.
How a mention is detected
We do not trust the engine's list of brands blindly. Each listed brand is checked against the answer text itself: a brand that is in the list but not in the answer is dropped, and positions are recalculated from where each brand first appears in the text. We would rather miss a mention than invent one.
Your brand counts as mentioned when one of its names or aliases appears as a whole word, ignoring case and accents. "Notion" matches "Notion AI" but not "Notional". Aliases shorter than 4 characters must match a brand name exactly, because short strings collide with ordinary words.
Tone is the engine's own label under a fixed rubric: positive when the brand is recommended or praised, neutral when it is listed or mentioned, negative when it is criticized or advised against.
Sampling, mention rate and confidence
Large language models are not deterministic: the same question can get a different answer minutes later. One answer per question is a sample of one, and we treat it that way.
By default, each scan records one answer per question and engine. Where multi-sampling is enabled for a plan, the same question is asked several times in the same scan. Extra samples are always fresh calls, never served from the cache, and the result becomes a fraction, for example named in 2 of 3 answers.
Mention rate = answers that named you ÷ all answers collected, using the latest answer for each tracked question in the last 30 days. Next to it we show a 95% confidence range (a Wilson score interval). With 12 mentions in 20 answers the rate is 60%, but the plausible range is about 39-78%. With 2 of 3 it is 67%, with a range of 21-94%: a small sample proves very little, and the range says so.
Samples of the same question are not fully independent, so read the range as a guide rather than a guarantee. When an engine showed no AI answer at all for a question, that question is left out of both the rate and the score: the engine did not pass over you, it did not answer.
The visibility score
Answer points = mention × tone weight × position weight
Engine score = 100 × (sum of answer points ÷ number of tracked questions)
Mention is 1 if the answer named you and 0 if it did not. With multi-sampling it is the share of samples that named you. When fewer than half the samples named you, tone and position are not recorded, so those answers get the neutral and "position unknown" weights.
| Tone | Weight |
|---|---|
| Positive (recommended, praised) | 1.00 |
| Neutral (listed, mentioned) | 0.80 |
| Negative (criticized, advised against) | 0.50 |
| No tone label | 0.80 |
| Position in the answer | Weight |
|---|---|
| 1st | 1.00 |
| 2nd | 0.96 |
| 3rd | 0.92 |
| 5th | 0.84 |
| 10th | 0.64 |
| 11th and later | 0.60 |
| Position unknown | 0.85 |
Each place after the first costs 0.04, down to a floor of 0.60. The penalty is deliberately gentle: being named at all matters far more than being named first.
Worked example. Three tracked questions. In the first, the engine names you first and recommends you: 1.00 points. In the second, you are listed third without praise: 0.80 × 0.92 = 0.736 points. In the third, you are not named: 0. Engine score = 100 × (1.00 + 0.736 + 0) ÷ 3 ≈ 58.
The score uses the latest successful answer per question from the last 30 days. An engine error never counts as "not mentioned": a failed call is ignored and the last good answer stands. A tracked question with no usable answer in the window counts as 0, and questions where the engine showed no AI answer at all are removed from the count.
Bands. 55 and above shows as Visible, 30 to 54 as Partial, and below 30 as Invisible.
Overall score. Each day, every engine gets a daily score from that day's answers, and the day's value is the average across the engines scanned that day. Your overall score is the average of the last 7 days that have data.
Windows and smoothing
- Engine scores use a rolling 30-day window: the latest answer per question. The change shown on an engine compares this window with the previous 30 days.
- The overall score is a 7-day average of daily values, so one unusual day moves it by about a seventh. Its change compares the last 7 days with the 7 before; until there are 8 days of history it shows as flat instead of an invented change.
- Trend charts plot one point per day (that day's latest answer per question) with no extra smoothing, so real day-to-day noise stays visible.
Freshness, cadence and the shared answer cache
Because prompts never name your brand, the same question asked of the same engine in the same country gets the same answer no matter who asks. Zene stores that answer once and shares it across accounts, keyed by question, engine and country. It keeps costs down, and it means two brands tracking the same question are measured against the same answer.
Cadence. Brands on a paid plan are scanned automatically every day. The Free plan scans ChatGPT and Gemini on demand, when you ask for it.
Daily is the target, not a promise for every engine: ChatGPT's answer is refreshed every three days, and on some plans other higher-cost engines may reuse a recent answer for a few days instead of asking again; refresh timing may be tuned further as we learn how often answers really change. If an account reaches its monthly scan budget, scans continue only on the lowest-cost engines and the others keep their last answers until the next period; at one and a half times the budget, scans pause until then. Every result records when its answer was produced, and that is the time Zene shows as the last scan.
Alert rules
- Two-scan confirmation. A "now appears on" or "dropped off" alert fires only when the change holds in your two latest full scans compared with the scan before them. A single scan that flips never alerts.
- Like-for-like questions. Only questions answered in all three of those scans are compared, so rotating question sets and newly added questions cannot create fake changes.
- Prompt-version guard. If our prompt for an engine changed between those scans, visibility and tone alerts for that engine are skipped until the scans share a version again. A method change is not a market change.
- Partial scans. Scans that covered only some questions (a targeted re-scan, or a scan cut short) never raise these alerts and are never used as the baseline.
- Multi-sampling threshold. With several samples per question, you count as visible on an engine when the named fractions across the compared questions add up to at least one half. One sample flipping from 0 to 1 of 3 does not cross that line.
- Tone alerts fire only for moves into or out of mostly negative mentions, with the same two-scan confirmation. Drift between positive and neutral is visible in the app but does not interrupt you.
- Score milestones (crossing 25, 50, 75 or 100 on the way up) only use scans that covered every tracked question.
Limitations
- Answers are not deterministic. The same question can get a different answer on the next run. Read trends and mention rates, not a single day.
- No personalization, no logged-in context. We ask as an anonymous API caller with no chat history, memory, custom instructions or account settings. A real person's answer can differ, sometimes a lot.
- The API is not the consumer app. The model and search behavior behind an API call can differ from what the ChatGPT, Gemini or Claude apps use at the same moment.
- Web search differs by engine. Gemini and Grok currently answer from model knowledge, so new content reaches them more slowly than it reaches ChatGPT, Claude and Perplexity.
- Location is approximate. Markets are country-level. City-level and in-app location signals are not reproduced.
- Brand detection has edge cases. Brands with generic-word names can be under- or over-counted, and a brand named only through a synonym we do not know is missed. Add aliases to reduce this.
- Search-page AI answers are not part of the score today. Answers such as Google AI Overviews only appear for some queries and in some countries. When a search surface shows no AI answer for a question, the right reading is "no answer", not "not mentioned", which is how we already treat no-answer results.
- Tone is a model label. Sentiment comes from the engine under a fixed rubric, not from a human reviewer.
- Scores are relative to your questions. Adding, removing or rewording questions changes the score. Compare like with like.
Public research data
Our public industry reports (Research and the live GEO for Industries database) use the same brand-blind prompt on ChatGPT, Gemini and Perplexity, with the same web search settings as customer scans: ChatGPT and Perplexity search the web, Gemini answers from model knowledge. They record the brands each engine reports naming, in order, and every source it cites; unlike customer scans, the reported list is used as is, without the answer-text check described above. An industry is only published once real collected data exists for it, and nothing is reordered by hand.
When the method changes
Prompts are versioned, and every scan records which version produced it. That is what lets the alert guard above tell a method change from a market change. Notable changes:
- September 2026. Answers are requested per market (country) instead of one default location, so brands outside the US may see a one-time shift. Mention rates with confidence ranges and multi-sampling support were added, and questions where an engine shows no AI answer were removed from scores.
Methodology questions
A 0-100 number per AI engine. Every tracked question earns up to 1 point from the engine's latest answer: 1 if your brand is named, reduced for later positions and for a neutral or negative tone, 0 if you are not named. The engine score is the average across your tracked questions. Your overall score averages the engines' daily scores over the last 7 days with data.
No. Prompts are brand-blind: the engine only gets the buyer's question, answers it, and then lists every brand it named. Mentioning your brand in the prompt would nudge the model to include you and inflate the very number being measured.
AI answers change from run to run, and your own chat includes memory, history, location and account settings that an API call does not have. Zene asks every engine the same way each time, so its numbers are comparable over time. A single manual check is not.
Brands on a paid plan are scanned automatically every day. The Free plan scans ChatGPT and Gemini on demand. ChatGPT's answer is refreshed every three days, some other higher-cost engines may reuse a recent answer for a few days instead of asking again, and every result shows when its answer was produced.
It is a 95% confidence range (a Wilson score interval). With only a few answers, the true rate could be well above or below what we observed, so we show the plausible range instead of treating a small sample as exact.
ChatGPT and Gemini, scanned on demand. The paid plans add Claude, Perplexity and Grok, with automatic daily scans.
No. A visibility alert needs the change to hold across your two latest full scans compared with the scan before them, on the same questions and the same prompt version. A single scan that flips never alerts.
See these numbers for your brand.
Run a free scan on ChatGPT and Gemini. No card needed, and every number comes with the method above.
