The AI Consensus Index

What the AI engines agree on, and where they do not.

Ask five AI engines for the best tool in a category and you tend to get five different answers. This page is a running record of what they actually recommend: which tools each engine names, and how far apart the engines are from each other. Every figure is a dated capture with the raw data attached. I do not tell anyone how to win AI search, and I do not decide the rankings. The engines produce the numbers. I record them and show the work.

Maintained by Vincent Wesley Couey, independent AI-citation analyst. Every figure below is traceable to an open dataset; the underlying data is CC BY 4.0 and DOI-backed where a permanent identifier exists.

0%
Full cross-engine agreement on the single top B2B tool (0 of 16 categories, 5 engines)
78
Divergence Score for AI video generators (of 100), across 5 engines
1 / 22
AI video tools named by all 5 engines in the recorded sample

Two separate studies, two separate engine panels, two dated snapshots. See each category and the provenance section below for the exact panel behind each number.

Categories currently measured (2)
Category measured

AI video generators

Question put to each engine: “What is the best AI video generation tool?

ChatGPTGeminiGoogle AI OverviewsGrokPerplexity
Divergence Score
78/ 100

For AI video tools, exactly one name survived all five engines: Creatify. Every other tool that cleared the bar was missed by at least one engine.

In the recorded 2026-06-17 sample, 1 of 22 tools (Creatify) was named by all 5 engines. Mean pairwise overlap of each engine’s top-5 was 0.223, a Divergence Score of 78 of 100.

Consensus Leaderboard: what the engines agree on
ToolNamed byTotal mentions
Creatify5 of 56
InVideo4 of 514
HeyGen4 of 513
Synthesia4 of 513
Google Veo4 of 510
CapCut4 of 56
Higgsfield4 of 56
Canva3 of 59
Runway3 of 57
Adobe Firefly3 of 56
Kling3 of 55
Fliki3 of 53
Seedance3 of 53

Tools named by at least 3 of the 5 engines in the recorded sample (13 of 22 tools cleared that bar). "Named by N of 5" counts how many distinct engines named the tool at least once; it is not an endorsement.

Method caveat · Per-engine naming counts are small and several tools tie at the top-5 boundary. Ties were broken by total cross-engine mentions, then tool name. Directional single-capture figure.

Captured · 2026-06-17
Sample · 44 recorded answers · 16 distinct queries · 22 distinct tools · 5 engines
Engine panel · Recorded live-engine outputs. ChatGPT (web), Gemini (web), Google AI Overviews (default Google Search), Grok (web), Perplexity (default web search).
Category measured

B2B / GTM software

Question put to each engine: “What is the best tool for [16 B2B / go-to-market software categories]?

ChatGPTPerplexityGeminiGroq (Llama 3.3 70B)Cohere (Command-A)
Divergence Score
61/ 100

Across sixteen B2B software categories, the five engines never once agreed on a single best tool. Zero categories out of sixteen.

In the recorded 2026-06-19 sample, 0 of 16 categories had all 5 engines name the same single top tool (0 percent full cross-engine agreement). Mean pairwise overlap of each engine’s top-5 was 0.387, a Divergence Score of 61 of 100.

Consensus Leaderboard: what the engines agree on
ToolNamed byTotal mentions
HubSpot5 of 537
Mailchimp5 of 515
Ahrefs5 of 514
SEMrush5 of 513
Salesforce5 of 512
ActiveCampaign5 of 511
Klaviyo5 of 511
Zendesk5 of 510
Calendly5 of 59
Jasper5 of 59

Tools named by all 5 engines somewhere across the 16-category audit (23 tools cleared that bar; the 10 most-mentioned are shown). Note the distinction the honesty floor forces: being named by all 5 engines somewhere is NOT the same as all 5 engines agreeing on the same top pick for one category, which happened in 0 of 16 categories.

Method caveat · Most tools appear as an engine top pick in only one category, so the per-engine top-5 sets rest heavily on tie-breaking and read as directional only. The robust, primary figure for this dataset is the recorded 0 percent full cross-engine agreement on the single top pick (0 of 16 categories). An earlier 3-engine GTM run (ChatGPT, Perplexity, Google AI Overviews) agreed fully on 36 percent of queries.

Captured · 2026-06-19
Sample · 716 recorded recommendations · 16 categories · 242 distinct tools · 5 engines
Engine panel · ChatGPT (web, GPT-5), Perplexity (web search), Gemini 2.5 Flash-Lite (grounded API), Llama 3.3 70B via Groq (model memory), Cohere Command-A (model memory). Two engines answer from training rather than live web and are labelled as model-memory.
The metric

How the Divergence Score is computed

For a category, we take each engine’s top-5 recommended tools from its recorded answers, then measure how much those shortlists overlap. The score is:

Divergence Score = 100 × (1 − mean pairwise Jaccard overlap
                          of each engine's top-5 shortlist)

0   = every engine returns an identical shortlist
100 = no two engines overlap on a single tool

The mean is taken over all engine pairs (10 pairs for a 5-engine panel). The score is directional, not a p-value: the shortlists come from a single dated capture, and where naming counts are small the top-5 boundary depends on tie-breaking. Each category above states its own caveat. The complementary positive figure, “named by N of the panel,” is the Consensus Leaderboard the same instrument produces.

Forward metric
Not yet measured

Volatility, coming next cycle

The same question, asked again, can return a different shortlist from the same engine. Volatility will measure how much a single engine disagrees with itself from run to run (the mean set-difference of its own top-5 across repeated captures). That is a separate thing from Divergence, which is the engines disagreeing with each other.

The datasets on this page are single captures, so Volatility has no honest value yet. It is deliberately shown here with no number. Volatility will be computed from repeated captures beginning the next measurement cycle, and reported per engine, per category, with the run count on its face. Until then, treat every ranking here as one dated reading, not a fixed property of any engine.

Method & provenance

Exactly what was measured, and how

AI video generators

Panel
ChatGPT · Gemini · Google AI Overviews · Grok · Perplexity
Captured
2026-06-17 · 44 recorded answers · 16 queries · 22 tools
Model note
Live-engine web captures. Counts are tool-naming frequency across recorded outputs; engines are reported per engine and never averaged into one score.
Honesty note
A logged-in ChatGPT session with prior site history is excluded from organic counts (personalization), consistent with the source dataset.

B2B / GTM software

Panel
ChatGPT (web, GPT-5) · Perplexity (web search) · Gemini 2.5 Flash-Lite (grounded API) · Llama 3.3 70B via Groq (model memory) · Cohere Command-A (model memory)
Captured
2026-06-19 · 716 recorded recommendations · 16 categories · 242 tools
Model note
Two engines answer from training rather than live search and are labelled model-memory; the grounded engine is sampled less densely, so the blended leaderboard leans slightly toward model memory and the per-engine comparison separates them. An earlier 3-engine run (ChatGPT, Perplexity, Google AI Overviews) is reported separately.
DOIs
10.5281/zenodo.20767878 (The AI Recommendation Audit 2026) · 10.5281/zenodo.20632768 (Who AI Recommends: GTM Tool and Source Citations (2026))
Honesty floor

Every figure is the recorded output of live AI engines on a dated snapshot, not a claim that any tool is objectively best and not a judgment of any company. Each panel is named per dataset; the two studies use different panels and are never merged into one blended “five engines.” “Named by N of the panel” is a checkable frame, not an endorsement. AI answers are volatile; re-run the published protocol to reproduce. Data is CC BY 4.0.

Reported, not ranked. Snapshots are dated and prior snapshots are not overwritten silently.