How AI engines choose which sources to cite
Four engines, four different citation habits. What we see in our own corpus of answers about UK businesses, and what it means for getting cited.
When an AI engine answers a commercial question, it usually shows its working: a list of sources beneath or beside the answer. Those citations are the most useful thing on the page, because they tell you exactly which documents the engine treated as authoritative for that question — and getting into that set is the work.
The engines do not behave alike.
They cite different kinds of source
Across our corpus of UK local-service answers, some engines lean on aggregators and directories, some on individual business sites, and some on editorial round-ups. For a customer, this is invisible. For anyone trying to get recommended, it changes the target entirely: getting listed on the right directory may matter more than anything you do to your own site, or barely matter at all, depending on which engine your customers use.
The only way to know which applies to you is to ask the engines about your service, in your area, and read the sources they return.
Some engines make their citations hard to read
Here is a technical finding worth publishing, because it affects any tool measuring this.
Google's Gemini does not, in our data, cite publishers directly. Every source entry we have recorded from it is a Vertex AI grounding redirect — a Google-hosted URL that stands in front of the real destination. Measured across our corpus on 31 August 2026 with our own citation-country-coverage command, that was 1,770 out of 1,770 Gemini source entries. Not most. All of them.
The practical consequence: any tool that reads the publisher's hostname straight out of the citation URL sees nothing at all from Gemini. Not "fewer citations" — zero, silently, with no error. In our own corpus Gemini represents roughly a quarter of everything we hold, so a naive implementation drops a quarter of the evidence and reports the remainder as if it were the whole picture.
We only found it because a number that should have been non-zero was exactly zero, which is the kind of result worth being suspicious of rather than relieved by. The fix is to recover the destination from the citation's own metadata; the general lesson is that "no data" and "zero" look identical in a dashboard and mean completely different things.
Citation is not recommendation
An engine can cite your page as a source for a fact while recommending three competitors for the job. Both are worth knowing, and they move independently, so they should never be collapsed into one number. Being cited is evidence the engine considers you credible. Being recommended is the thing that pays.
Cited sources are not always local
Because these engines assemble answers from a broad index, cited sources for a UK query are not automatically UK businesses. We have measured cross-continent citation directly: a firm in Birmingham, Alabama appearing among the sources for a question explicitly about a UK city.
This matters if you are using citations as a target list. A source that cannot serve your customers is not an opportunity, and a list that quietly includes them overstates how much work is available. Our own outreach lists withhold any host we have not positively established serves the UK, and count what was withheld on screen rather than dropping it silently — an honest "we're not sure about 12 of these" beats a confident list with foreign firms hidden in it.
What to do with this
- Get the citation list for your own questions. Not a generic list of directories — the sources actually returned for what your customers ask.
- Sort by how often each source appears. A source cited across many answers is worth more than one that appeared once.
- Pursue the ones you can realistically get into. A listing, a review presence, a mention in a round-up, a case study on a supplier's site.
- Check they are real targets. A cited source that serves a different country is not a lead.
- Re-measure. The set changes.
Every check we run keeps the full source list behind every answer, so the target list is evidence rather than opinion.