Our own measurement, published so it can be checked
What AI searches for when someone asks it to recommend a business
When an engine answers “who should I call”, it does not search for what the person typed. It writes its own searches first. We kept 595 of them out of 188 stored answers from ChatGPT, Gemini and Perplexity. They are the step almost nobody publishes, and they say more about why a business is missing than the answer does.
ChatGPT ran 478 searches of its own across 78 answers.
That is 6.1 searches for every one question it was asked, and 12 in its busiest single answer. Gemini ran 117 across its own 78 answers, 1.5 each, and never more than 2. Perplexity exposed none of its searching in 32 answers.
So a business is not competing to match the buyer’s question. It is competing to match whatever the engine decided to go and look for, which on this sample was shorter and more specific than the question itself: a median of 8 words against the buyer’s 12.
Source. 595 searches recovered from 188 answers to 48 question wordings · ChatGPT (gpt-5.4), Gemini (gemini-3.6-flash), Perplexity (sonar), each through its API with web search on, not the consumer apps · sampled 24 and 25 August and 16 September 2026 · 142 of the 188 answers exposed at least one search · recomputed 1 October 2026 from the stored raw API responses.
How much each engine searched
A “search” here is one query string the engine issued while answering, as the provider returned it with the response. We did not infer these; they are stored strings.
| Engine | Answers | Answers exposing a search | Searches | Per answer | Most in one answer |
|---|---|---|---|---|---|
| ChatGPT (gpt-5.4) | 78 | 75 | 478 | 6.1 | 12 |
| Gemini (gemini-3.6-flash) | 78 | 67 | 117 | 1.5 | 2 |
| Perplexity (sonar) | 32 | 0 | 0 | 0.0 | 0 |
| All three | 188 | 142 | 595 | 3.2 | 12 |
Perplexity’s zero is a limit of what its API returns, not a finding that it does not search. It cites sources in its answers, so it plainly retrieves. We cannot see the queries, so Perplexity is absent from everything below.
Source. 01_Intel/research/icp_test.json (128 answers) and 05_Engineering/work/runs/cinc_panel.json (60 answers), Vector Index’s own research runs · no client data · 583 of the 595 search strings are distinct.
157 of ChatGPT’s 478 searches named one site to check
26.4% of the searches used the site: operator, which restricts a search to a single domain. They targeted 97 distinct domains across 38 answers. Every one came from ChatGPT; Gemini used it in none of its 117.
Two patterns run through them, and both are verbatim below. In the first, the engine has a company name already and goes to a third party to check it:
site:bbb.org Cincinnati HVAC BBB Corcoran Harnist Heating Airsite:bbb.org Cincinnati HVAC BBB JonLe Heating Coolingsite:bbb.org Cincinnati HVAC BBB Arlinghaus Heating Airsite:napfa.org Cincinnati OH fee-only advisor truepoint wealth counselsite:ohiobar.org certified specialist family relations law Cincinnati attorney
In the second, it goes to the company’s own domain looking for particular words:
site:plasticsurgerygroup.net Cincinnati plastic surgery group official board certifiedsite:theplasticsurgerycenterofnorthernkentucky.com Cincinnati plastic surgeon official board certifiedsite:sportysacademy.com Cincinnati flight training official Sporty’s Academy Clermont Countysite:trihealth.com Cincinnati cosmetic plastic surgery official board certified
The word “official” appears in 162 of the 595 searches, 27.2%, and 28 of those also use site:. The most-targeted domains were bbb.org (13 searches), plasticsurgery.org (9) and abplasticsurgery.org (7).
What this does not show: that a page satisfying one of these searches changes the answer. We recorded the search, not its effect. The honest reading is narrower and still useful. On this sample the engine was running named checks against specific sites, so the question “does our own site state plainly what we are and what we are certified in” is answerable rather than rhetorical.
Source. 157 of 595 search strings contain the substring site: · domains parsed from the operator · 38 of the 188 answers · all ChatGPT.
What words the searches contained
Each row counts search strings containing the substring, pooled across engines. These are string matches, which is all they are: a search containing “bbb” is not proof the engine read a BBB page.
| Substring | Searches | Share of 595 | ChatGPT, of 478 | Gemini, of 117 |
|---|---|---|---|---|
site: | 157 | 26.4% | 157 | 0 |
| official | 162 | 27.2% | 162 | 0 |
| best | 90 | 15.1% | 43 | 47 |
| review | 36 | 6.1% | 35 | 1 |
| bbb | 30 | 5.0% | 30 | 0 |
| cost, price, pricing, how much | 17 | 2.9% | 13 | 4 |
| licens | 15 | 2.5% | 14 | 1 |
| yelp | 2 | 0.3% | 2 | 0 |
| 1 | 0.2% | 1 | 0 | |
| near me | 0 | 0.0% | 0 | 0 |
“Near me” appears in none of the 595. The engines rewrote proximity into a place name instead, which is the opposite of how the phrase is usually discussed.
The two engines behave differently enough that one test tells you little
47 of Gemini’s 117 searches contain “best”, 40.2%, against 43 of ChatGPT’s 478, 9.0%. Gemini mostly restated the question as a category-and-city query. Its searches, verbatim:
fiduciary fee only financial planner Cincinnati Ohiobest fee only financial advisor Cincinnati OHtop wealth management financial advisors Cincinnatibest cosmetic surgery centers Cincinnati Ohio
ChatGPT, on the same kind of question, went from the category to named companies and then checked them one at a time. Reading only one engine would leave a business optimising for the wrong step: a category-and-city page for Gemini, a checkable record on third-party sites for ChatGPT.
Source. 117 Gemini and 478 ChatGPT search strings · the “best” counts are substring matches on those strings.
How we counted
- What a search is. One query string the engine issued while producing the answer, returned by the provider alongside the response and stored with it. No search was reconstructed, guessed or parsed out of prose.
- Engines. ChatGPT (gpt-5.4), Gemini (gemini-3.6-flash) and Perplexity (sonar), each through its API with web search enabled. API answers are not consumer-app answers: system prompt, retrieval and routing all differ.
- Location. None was set, and no account was signed in. Where a city appears, it is in the wording of the question only, which is not evidence of where a user is.
- Substring counts. Matched case-insensitively against the search string. A search counts once per row however many times the substring occurs.
- Distinctness. 595 searches, 583 of them distinct strings. Repeats across answers are counted each time, because each one was issued.
- Questions. 48 distinct wordings. The 6 in cinc_panel.json are a subset of the 48 in icp_test.json, checked rather than assumed, so 48 is the union and not a sum. 45 of the 48 produced at least one search.
- Client data. Excluded by construction: only the two research files above were read, neither of which is under the client folder, and every search string was additionally screened against every client name and domain on file. 0 were screened out, and the screen was tested by running it with two substrings that do occur, which removed 478 of the 595.
Limits
- Perplexity is missing from the findings. Its API exposed no searches in 32 answers. That is a gap in what we can see, not a finding about the engine.
- What the provider chooses to show us. These are the searches each API reported. An engine may run searches it does not report, and 46 of the 188 answers exposed none at all. Nothing here is the complete retrieval behaviour of any engine.
- One city, mostly. Most questions name Cincinnati. 10 of the 48 name no city, such as “best executive outplacement firms”, and those produced 63 of the 595 searches. Nothing here says what happens in another metro.
- Dated. 24 and 25 August and 16 September 2026, on the model versions named. Retrieval behaviour is a product decision and can change without notice.
- Searches, not effects. We recorded what the engines looked for. We did not test whether satisfying any of those searches changes an answer, and nothing on this page is a causal result.
- The questions were ours. We wrote them for our own research. They are not drawn from search logs and are not a sample of what buyers ask.
- Not measured: how many businesses each answer named. The stored research file carries a per-answer list of firm names, and it is not fit for that purpose, so we have not published a count. For one home-services answer it holds 171 entries including a concrete contractor and an electrician, neither of which appears in an answer about replacing a furnace. It is a harvest from the retrieval layer, not from the answer. Counting named businesses properly needs an extractor whose output has been read against every answer, and this pass did not do that.
Every search we recovered
All 595, one per row, with the question that produced it, the engine, the model, the date and the substring flags used above. The counts on this page can be reproduced from this file alone.
Download the data (CSV, 595 rows)
Some rows name real local businesses, because the engine named them in its own search. Their presence says only that: the engine searched for them on the date shown. It is not a statement by us about any of them.
Which sites AI engines cite · How we measure · Why ChatGPT does not mention a business · Our scope and prices
Questions we get asked about this.
What is query fan-out in AI search?
It is the set of searches an engine writes for itself when answering a question, instead of searching for what the person typed. In 188 stored answers we recovered 595 of them. ChatGPT averaged 6.1 searches per answer and Gemini 1.5.
Does ChatGPT search for my business by name?
On this sample it did for some businesses. 157 of ChatGPT's 478 searches used the site: operator to check one domain, and several named a specific local company and looked it up on bbb.org or a professional directory. We measured the searches, not whether what they found changed the answer.
Why does the engine search for something different from the question?
Because it reformulates first. On this sample the engines' searches were shorter and more specific than the questions, a median of 8 words against 12, and often replaced a phrase like "near me" with an actual place name. None of the 595 searches contained "near me".
Do all the AI engines behave the same way?
No, and that is the practical finding. 40.2% of Gemini's searches contained "best" and it mostly restated the question as a category and a city. ChatGPT went from the category to named companies and then checked them individually. A business that tests one engine sees one of those patterns.
Can I see the data behind this page?
Yes. All 595 searches are published as a CSV with the question, engine, model and date for each, so every count on the page can be recomputed from it.
Does this prove what to change on my website?
No. It records what the engines looked for, not what works. We did not test whether satisfying any of these searches changes an answer, so nothing here is a causal result. What it does give you is a specific thing to check rather than a guess.