How to Set Up AI Visibility Tracking
Setting up visibility tracking means building a question set, freezing it, asking it across both channels on a schedule, and keeping every run in full. The readings are whether a company was named, how it was described, who was named instead, and which sources the answer was built from. None of those is a number a rank tracker can produce.
What is being measured, exactly?
An answer, not a position. The reading a run produces is whether a company was named in response to a question, how it was described, which companies were named alongside it, and which sources the answer was built from. Nothing in that list is a number a rank tracker can produce, because there is no position to record. An answer either names a company or it does not, and the words it uses matter as much as the mention.
How is a question set built?
From the questions a buyer actually asks, in the words they would use, plus the category terms an industry uses about itself. Both belong, because they return different sets of names. Falkview's own first run recorded the category question and the buyer's question producing almost entirely different answers, which is a reason to ask both rather than assume one predicts the other.
Coverage across the buying sequence is the second rule. Questions that establish a category, questions that compare options, questions about how the work is done, and questions asked in the last hour before a decision all behave differently. A set weighted entirely toward one stage measures one moment of a decision and reports it as visibility.
Freezing the set is what makes it a measurement rather than a survey. Questions are fixed before the first run, and a question added later starts its own series rather than joining the existing one. A set that drifts produces a report whose second month cannot be compared with its first, which is an expensive way to learn nothing.
Size is the last decision, and a smaller set asked weekly is worth more than a large one asked once. Every question added multiplies the work of every future run, and a set nobody can sustain quietly becomes a set that gets sampled rather than asked, at which point the series has stopped comparing like with like.
Which surfaces belong in a run?
Both channels, kept apart in the record. The AI features inside search results and the standalone assistants people open directly are different products with different retrieval behavior. Collapsing them into one figure hides the most awkward finding available: ranking well on a query whose composed answer recommends a competitor.
A surface that could not be reached on a given day is recorded as not measured. Recording it as an absence would be a different claim, and confusing the two is the most likely way for a report to be confidently wrong about a company. Falkview's own first run made a version of that mistake, published the correction rather than removing it, and turned it into a written rule in the procedure.
What is recorded from each run?
The full answer text, rather than a summary of it. A summary loses the wording, and the wording is often the finding. A company named as a specialist in one thing and a company named as one option among several have both been mentioned, and the two mentions are not worth the same to whoever is reading.
Alongside the text: the sources cited, the other companies named, the date and time, and the conditions the run was taken under. A position varies with location, device and personalization. An answer moves between runs, and moves again when its provider updates the system behind it. Strip the conditions off a reading and there is nothing left to compare the next one against.
How is a series read?
By direction over a period, never by the value of a single run. Answers are probabilistic, so two consecutive runs can differ without anything having changed on the site or off it. What carries information is a movement that persists across several runs, read against the record of what went live between them.
Citations are the line that converts into work. Where the sources behind a subject's answers are a publication, a directory entry and a competitor's comparison page, those are three specific things to address. Mention rate says a company is losing. The source list says what it is losing to, and two of the three are usually not on the company's own site.
What can a measurement not tell you?
Why anything moved. An answer that changes between runs may reflect work done on the site, a newly earned independent source, or a provider updating its system that week. Separating those from outside is rarely possible, and a report assigning a cause it cannot support is worth less than one recording the sequence and saying plainly which part is inference.
Nor does a run predict the next one. A single reading samples a system that varies, which is why a screenshot of an assistant naming a company is not evidence of a state, and a screenshot of one omitting a company is not evidence of a problem. Both are moments, and a moment is not a measurement.
Where is the procedure published?
Falkview publishes its measurement procedure under how we measure, including the conditions each run records and the distinction it holds between an observed position and prompt-level visibility. What Our Own Baseline Said on Day One is the first run done under that procedure, published in full, including the conclusion the firm got wrong and corrected in public.
What AI Search Visibility Is, and How to Measure It covers what the readings mean and why rank tracking cannot produce them. How to Monitor Your Brand in AI Search covers the other half of the same work: asking about a company by name rather than asking the category question. The standing version of both is sold as AI visibility monitoring.