// Methodology · Answer Monitor
Most AI visibility scores are one coin flip
Ask an engine the same question five times and get five answers. So we run it like a poll, and every finding ships with its run count.
Updated 1 September 2026
// The problem
Most tracking measures noise
Identical prompts return different citations run to run, and engines swap most of their cited sources weekly. A monthly rank report is last month's noise presented as this month's trend.
// The standard
Seven rules we don't break
Sample like a poll
The same prompt can return a different answer on the next run, so one reading is a guess dressed up as a fact. We treat every prompt the way a pollster treats a question: run it enough times to see the real spread of answers, then report that spread instead of a single draw.
Report each engine on its own
ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews and Google AI Mode disagree with each other constantly, so we run all six and never merge the results. A blended score hides the one thing that matters, which is which engine said what, for which prompt. Every finding names its engine.
Run every prompt several times
One run tells you what an engine said once. It does not tell you what an engine usually says. So every prompt runs on every engine several times, and a finding seen once is never merged with one that repeated.
Show the run count on every finding
Cited in one run out of ten and cited in ten out of ten are both true and completely different facts. Every finding on every report carries both numbers, so you can tell a stable answer from sampling noise without asking us to check.
Split prompts by persona
A first-time buyer and a technical evaluator type different questions into the same engine and get different answers back. Writing one prompt set for an imaginary average customer erases that difference before you ever see it. Our prompt sets are split by persona from the start.
Start a new series when the instrument changes
Adding an engine, or rewording a question, changes what the number means. Comparing a reading taken after that change to one taken before it treats two different instruments as one continuous trend. So when the instrument changes we start a new series, label it, and let anyone already mid-quarter finish on the old one before we switch them.
Track five-stage conversations
Buyers ask a question, then a follow-up, then another, before an engine ever names a brand. A single prompt only ever catches one link in that chain. We run each persona through five connected stages and record where you enter the answer and where you fall out of it.
// How we prompt
We track the whole buyer journey, as a conversation
One-shot prompts miss how recommendations form. We run each persona through five connected stages and watch where you enter the answer, and where you fall out.
Problem
"Why do my crypto transactions keep failing?" The moment before your category is even on the table.
Exploration
"What are the best options for X?" Where the model builds its shortlist, and where category ownership is won or lost.
Comparison
"X vs Y": head-to-head, where sentiment and framing decide the winner.
Validation
"Is X legit / audited / safe?" Where a wrong or stale fact quietly kills the deal.
Selection
"How do I get started with X?" Where the model either hands the customer to you, or to a competitor.
// What lands in your report
Four numbers, reported separately
Per engine, per persona, per journey stage: how often you appear, where you rank in the shortlist, whether you are recommended ahead of competitors, who you are cited alongside, the sentiment and framing around you, and whether the facts the model states are true.
Mention is whether your brand name appears in the answer text. Citation is whether your page gets linked as a source. Recommendation is whether the engine tells the buyer to choose you, not just discusses you. Accuracy is whether what the engine says about you is true. A brand can score high on one and low on another, and averaging them into a single number would hide exactly that gap. So we report all four separately, every time.
Measurement you can defend to your board
Answer Monitor applies this standard to crypto and fintech brands first. Get the Diagnostic and we'll baseline your category with the run counts shown.
Get the Diagnostic →Method FAQ
How many engines does this cover?
Six, every time: ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews and Google AI Mode. Each one is reported on its own line. We never blend them into a single average.
Why does one run not count as a finding?
The same prompt can come back with a different answer on the next attempt, so a single run is a coin flip. We only call something a finding once it has survived several runs on the same engine.
What does a run count tell me?
It tells you how stable the answer is. Cited in one run out of ten and cited in ten out of ten are both technically "cited," but they are different facts about how reliably that citation happens. We show both numbers on every finding rather than collapsing them into one.
Why split prompts by persona instead of using one set of questions?
A first-time buyer and a technical evaluator ask different questions and get different answers from the same engine. A single generic prompt set would average those two audiences into a customer who does not exist.
What happens when you add a new engine or change a prompt?
We start a new series and label it. Comparing a reading taken after the change to one taken before it would treat two different instruments as one continuous trend, so we never do that. Anyone already mid-quarter on the old series finishes it before moving to the new one.
Why measure a five-stage journey instead of a single prompt?
Buyers rarely get a recommendation from one question. They ask, get a shortlist, compare options, check credibility, then ask how to start. A single prompt only ever catches one of those five stages, so we run each persona through all five and record where you enter the answer and where you drop out.
Our approach is informed by published research on AI-answer variance, notably Kevin Indig's Growth Memo and its data partners (AirOps, SISTRIX). The measurement standard and its application are our own.