SERPInsight records Google AI Overviews for tracked search queries on a schedule, and publishes what it finds. This page states how that is done, what the figures mean, and — as importantly — what they do not cover. Every data page on this site links here.

How is the data collected?

Tracked queries are submitted to a search-results API on a fixed schedule. Each response is archived verbatim before anything parses it, then read for AI Overview presence, the domains cited, the organic results, and other SERP features. Every run is recorded whether it succeeded or failed; a failed run is data too, and is never quietly dropped.

Archiving the raw response first is deliberate. Google changes AI Overview markup without notice, and the archive is what allows history to be re-read against an improved parser instead of being lost.

How often is each query checked?

Queries sit in one of four refresh tiers. High-demand queries are checked daily, moderate ones weekly, the long tail monthly, and queries dormant for twelve months quarterly before retirement. Historical data is kept when a query retires; only the ongoing cost stops.

What does Source Overlap mean?

Source Overlap is the share of domains cited in an AI Overview that also appear in the organic top 10 for the same query, in the same check. An overlap of 100% means the AI Overview cited only pages that already rank. A low overlap means it drew on sources that do not rank, which is where citation opportunity exists that organic position would not reveal.

The term is used consistently across this site. It is never called citation overlap or organic match rate.

What does “cited without an organic ranking” mean?

It counts citations that occurred on queries where the domain did not appear in the organic top 10 at the time of the check. Expressed as a percentage of that domain’s recorded citations, it answers a question conventional rank tracking cannot: how much of a domain’s answer-layer visibility is independent of where it ranks.

When does a page get published?

A query page is published only when all of the following are true. The threshold is not relaxed to increase the number of pages.

Domain profiles have their own floor: a domain must be cited across several distinct queries before a profile publishes. A domain cited once is a single observation, not a pattern, and a page built on it would be misleading.

Pages below either threshold are marked noindex, excluded from sitemaps, and say plainly on the page that there is not enough data yet.

What do the sample sizes mean?

Every aggregate figure on this site carries the number of observations behind it. A citation rate of 40% across 5 queries and the same rate across 5,000 are very different claims, and the site never presents them as equivalent. If a figure is quoted from this site, the sample size and the as-of date should be quoted with it.

What this data does not tell you

How may this data be used?

Figures on this site may be quoted freely with attribution, provided the sample size and as-of date are carried with them. The requested form is:

SERPInsight, AI Overview Citation Index, accessed [date]. [page URL]

How are errors and removals handled?

Every data page carries a control to report an error. Domain owners may request removal of their citation profile, which is honoured without argument and takes effect immediately — the page is deindexed and the mirror row marked accordingly.

Query pages involving personal names, contact details, or any sexual content involving minors are blocked from publication by automated screening, with borderline cases held for human review rather than published.