HIGH SALIENCE / RESEARCH / AI SEARCH BENCHMARK 01

STATUS: IN PROGRESS · METHODOLOGY PUBLISHED AHEAD OF RESULTS

AI Search Benchmark 01

1,000 commercial SaaS queries. Three AI surfaces. One question: who gets recommended, and why?

This page publishes the methodology before data collection: the claim under test, the exact method and what we'll report. Publishing it first is deliberate: it means the findings can't be quietly shaped after the fact. Findings appear here when the benchmark has actually run.

01The Claim We're Testing

"AI answers are just repackaged Google results."

If that claim were true, the brands recommended by AI surfaces would closely track classic organic rankings, and the sources cited would be the same pages that rank. This benchmark measures how true that actually is on commercial, money-intent queries, where the stakes are highest and the training-data shortcuts weakest.

The answer decides how much of the game is still classic SEO and how much is genuinely new work: the degree to which answer engine optimization and generative engine optimization are separate disciplines rather than rebranded ones.

02Methodology

Fixed queries, logged answers, counted citations.

METHODOLOGY v1.0 · PUBLISHED AUGUST 29, 2026. AMENDMENTS, IF ANY, WILL BE VERSIONED AND DATED HERE.

  • Query set: 1,000 commercial SaaS queries across 20 categories, split across category, comparison, alternative and recommendation intent. The full list ships with the dataset.
  • Surfaces: ChatGPT (with browsing), Gemini and Google AI Mode, sampled in the same window from clean, logged-out sessions with memory and personalization disabled.
  • Sampling conditions: the exact collection dates, country/location, language, and the model or product version where observable are recorded with every run and published with the dataset.
  • What we log: every brand named, its position in the answer, sentiment of the mention, every cited URL and domain, and the classic Google top-10 for the same query, collected in the same window, as a control.
  • Runs: each query sampled multiple times per surface to account for answer variance; the repetition count per query per surface is fixed before collection starts and reported with the results. We report frequency, not single answers.
  • Coding rules: what counts as a "recommendation" versus a passing mention, and how sentiment is classified, are defined in a written codebook that ships with the study, written before collection, not after.
  • Citations: deduplicated to the canonical domain; per-URL detail preserved in the raw data.
  • Failures: refusals, errors and non-answers are logged and reported, not silently dropped.
  • Dataset: the raw query-level data ships with the published study so the findings can be checked, not just believed.

03What We'll Report

Findings publish when the benchmark has run. Not before.

The published study will report, at minimum:

  • How often recommendation-intent answers name brands that do not rank in Google's top 10 for the same query.
  • How concentrated citation influence is: how few domains account for the majority of citations across the three surfaces, and how much of that influence is third-party.
  • How often the three surfaces agree on the top recommendation, and how often they contradict each other.
  • Per-category splits, full result tables and the citation census, with a downloadable dataset.

We publish numbers when we have them and not a moment earlier. A benchmark announced with invented findings would be exactly the kind of marketing this research program exists to replace.

The same measurement discipline runs inside every client engagement: it's how our GEO program and AI Search engagements get scored, and the citation census feeds the source maps behind Source Authority campaigns.

04Limitations

What this study won't be able to tell you.

AI answers vary by session, region, account state and time; frequencies will describe the sampling window, not a permanent state. The query set is SaaS-weighted, and other categories may behave differently. Citation visibility differs by surface, so cross-surface citation comparisons are directional. We publish limitations so the findings can be trusted; a benchmark without them is marketing.