How AI Search Engines Choose Their Sources
No published rule exists for how an answer-composing system decides that one source is worth quoting and another is not. Retrieval is described publicly by at least one provider; selection among the retrieved pages is not described by any of them. What can be recorded instead is which sources the answers on a subject actually cite, across a fixed question set run repeatedly, and that record is enough to work from.
Is there a published rule?
No published rule exists for how an answer-composing system decides that one source is worth quoting and another is not. The companies building these systems have not documented the weighting, and there is no specification an outside party can check a claim against. An article naming the factors in order, with weights attached to each, has invented them.
Two claims get conflated constantly here, and separating them is most of the value in this piece. How pages are retrieved is described publicly, in general terms, by at least one provider. Which of the retrieved pages ends up cited, and why one is preferred over another, is not. Confidence about the first is not a license for confidence about the second.
What is documented about retrieval?
Google's developer documentation describes its AI-generated search answers as being drafted from a set of pages a conventional search retrieved first, a step the documentation calls grounding. Two consequences follow from that description without any inference about weighting. A page ordinary search cannot reach is not a candidate. And an answer is only as complete as the specific pages the retrieval step happened to pull that day.
Whether other providers' systems work the same way is not something an outside party can assert, and this piece does not assert it. Each provider documents its own products, and the descriptions do not all exist at the same level of detail. Treating one company's published account as a general law of the category is a small piece of invention that spreads easily.
What can be observed instead?
The citations themselves. Where an answer shows the sources it used, that list is readable by anyone asking the question, and it can be written down. Across a frozen question set, asked repeatedly, the recurring sources become visible: the publications, directories, reference sites and competitor pages that answers on a subject are built from.
One caution belongs with that record. A citation list shows what an answer drew on for that run, on that day, and answers vary between runs, so a source appearing once is not yet a pattern. What makes the list usable is repetition: the sources that keep returning across a frozen question set are the ones worth acting on, and the ones appearing once are worth noting and nothing more.
A recurring source list is evidence about a subject rather than a rule about a system. Knowing that answers on a category are consistently drawn from a particular publication does not explain why that publication is preferred. Knowing it is still enough to act on, because the publication is a specific, reachable thing and the weighting behind it is not.
What does the pattern suggest, and how firmly?
Falkview's working expectation, stated as an expectation rather than as a finding, is that a page stating a plain, checkable fact in a fixed structure is easier to draw from than a page of adjectives. Directory and register entries, reference pages and editorial coverage naming specifics all share that shape. A page asserting excellence without stating anything checkable gives a system nothing to quote.
A second expectation, held the same way, is that agreement between independent sources matters. Where several independent accounts of a company agree, a system has firmer ground to stand on than where they contradict each other. No provider has published a weighting confirming this, and this firm has not measured a population large enough to claim it. Both expectations are written down so a later run can score them rather than reinterpret them.
Why does the distinction matter commercially?
Because the work does not depend on knowing the rule. A citation analysis produces a list of specific pages that answers on a subject are currently built from. Some of those pages can be corrected, some can be earned, and some belong to competitors and cannot be touched. Every one of those decisions is available without a mechanism behind it.
Claiming the rule creates a debt instead. A plan justified by an invented weighting has to keep explaining the weighting each time a result arrives that does not match it. A plan justified by an observed citation list only has to keep observing, and the observations accumulate into something a client can read for themselves.
What does this mean for the third-party half of the work?
Coverage is chosen against evidence rather than against a media list. Where answers on a subject already draw from a particular set of publications and reference sites, those are the places worth pitching and the entries worth correcting. Where they draw from a competitor's own comparison page, the useful response is a page of the company's own stating the same facts as plainly.
Accuracy is the standard rather than volume. A large set of low-quality placements produces descriptions that disagree with each other, which makes a company harder to reconcile rather than easier. Falkview records new placements and citation lists separately, and says which is an observation and which is an attribution, so a client can tell the two apart in a report.
Which pieces carry this further?
What Is Digital PR, and Why It Now Shapes AI Answers covers the discipline behind the third-party half, including how coverage is chosen and what it is measured against. Why Your Competitor Gets Cited and You Don't covers the same question from the other end: what a cited page usually did that an uncited one did not.
How Does ChatGPT Decide Which Companies to Recommend? applies the same rule about undocumented mechanisms to the company-selection question rather than the source-selection one. Falkview scopes the outreach and coverage work as digital authority.