SeoVision

Data Sources & Methodology

Last updated September 11, 2026

A research tool should be transparent about where its numbers come from and how confident they are. This page explains exactly that. The short version: some of what SeoVision shows is observed directly — we ask AI assistants and read the answers they actually give, we crawl your site, and we read your own connected accounts. Some of it is a reading of those observations. And some is estimated from licensed third-party data. All three are useful; they are not equally precise, and the rest of this page says which is which.

Where the data comes from

  • AI-assistant answers we query directly. For AI Visibility we send your tracked prompts to the assistants themselves and read the answers they return. We call each model through its own interface, with live web search enabled where the assistant supports it, so what we store is the answer that was actually given — observed, not reconstructed from anonymized browsing panels.
  • Our own models, reading those answers. An answer is a paragraph of prose; a dashboard needs fields. So a second model reads each answer and pulls out the structured part: whether your brand is mentioned, how prominently, the sentiment around it, and which of the other names in the answer are genuinely your competitors rather than publishers or tools. Links and domains that appear in the text are extracted by exact parsing, not by a model. A model also scores how hard each cited domain is to be quoted by — that is what the Difficulty column on your Prompts dashboard reads. These fields are our reading of an observed answer: far better than a keyword match, and still a reading.
  • Your connected accounts. When you connect Google Search Console we read your own first-party search data — impressions, clicks and positions. When you connect a store or CMS we read and publish your site content. Both are your data, as the platform itself reports it.
  • Licensed SEO data. Search volume, keyword difficulty, domain rating, referring domains, backlink totals and search-results pages come from established third-party SEO data providers. These are provider estimates, produced by their crawls and their models rather than ours.
  • Our own site crawl. The SEO audit crawls the page or site you enter and evaluates technical, on-page, content and structured-data signals directly from the live HTML.

We don't publish the exact list of providers, feeds and models behind every single figure — that combination is our own work — but we do commit to being clear about which of the three kinds below a number is, and where its limits are.

How AI Visibility is measured

This is the most important thing to understand, and it is where SeoVision differs from tools that estimate AI traffic from anonymized browsing panels. We do not guess which answer an assistant gave — we ask it. For each tracked prompt we call the assistant, store the response, and then read it for whether your brand and your competitors appear, how prominently, with what sentiment, and which sources are cited.

Those are two different steps and they deserve different amounts of trust. The answer itself is observed: it is a real response to a real prompt at a real moment, and we keep the text so you can read it yourself. What we extract from it — mentioned or not, how prominent, positive or negative, competitor or not — is a readingof that text, done by a model working to a fixed schema. On top of that sits generalisation: an assistant's answers vary between runs and over time, so a single scan is one sample, and roll-ups across many prompts show direction and relative size rather than an exact share.

SEO audit & scores

The audit combines what we crawl from your live site with licensed SEO data and, when connected, your Google Search Console figures. Individual checks are measured against the page; scores, grades and roll-ups are computed from those checks and are meant for relative comparison and tracking improvement over time, not as an absolute measured count.

Three kinds of number

These three words describe the kinds of figure this product works with. They are how we talk about our own data on this page — not badges you will find printed beside every number in the dashboard.

  • Observed. Read directly off something real: an assistant's actual answer, a Search Console figure, a check read straight from your page, a link found in the HTML.
  • Estimated. Produced by a model or a dataset rather than counted. That covers third-party figures — most search volume, keyword difficulty and domain rating — and our own models' reading of an answer: mention, prominence, sentiment, competitor classification, domain difficulty.
  • Computed. Derived from other numbers: an audit score, a share-of-voice roll-up across prompts, or estimated traffic — search volume multiplied by the click-through rate typical of the position you hold. Good for direction and for watching change, not an absolute count.

If a figure anywhere in SeoVision ever reads as more precise than the list above says it is, we would rather hear about it than have you discount everything else — tell us and we will fix it or explain it.

Cadence & coverage

  • Refresh: AI Visibility scans run on a regular schedule (typically daily); search-volume and other market metrics refresh less often (weekly to monthly) because the underlying datasets update less often.
  • AI engines covered: ChatGPT, Claude, Gemini, Perplexity, Grok, Copilot and DeepSeek, plus Google AI Overviews and AI Mode tracked through the search results page.
  • Geographic & language coverage: follows the markets and languages you configure in your project, within the limits of your plan.
  • Audit & content: the SEO audit works on any public URL; the content engine publishes to Shopify, WordPress, WooCommerce, WordPress.com, Wix, Webflow and Next.js, and to any other site through our publishing API.

Limitations

  • Estimated figures carry uncertainty and are best used for relative comparison and trend direction, not exact accounting.
  • Assistant answers can change between runs and over time; a scan is a sample of what was returned at that moment.
  • Fields we read out of an answer — mention, prominence, sentiment, competitor classification, domain difficulty — are a model's interpretation of observed text. They are consistent and you can always open the answer they came from, but an unusual answer can be read wrongly.
  • Coverage is limited to the engines, markets and languages configured for your project; absence from our data is not proof of absence in the market.
  • Third-party data coverage and attribution models can change, which can shift figures over time.