The Hidden Bias in Your AI Visibility Score: A Pharma GEO Guide
A brand manager tells the steering committee their product is cited in 40% of relevant AI answers. A colleague, on a different product, reports 10%. Leadership concludes one team has mastered Generative Engine Optimization (GEO) and the other has not. That conclusion is almost certainly wrong.
Citation rate measures the question pool, not the brand
GEO citation rate is not an intrinsic property of a brand or a piece of content. It is the output of a specific instrument: the pool of questions used to query the generative engines, and the sample of engines and prompts run against it. The foundational GEO research from Aggarwal showed that specific content tactics — adding statistics, citing authoritative sources, quoting experts — can shift citation prominence by tens of percentage points depending on the query set tested1. If tactics alone can move the number by 40 points, the composition of the question pool moves it by even more. A pool built from ten narrow, branded, low-competition questions will naturally return a high citation rate. A pool built from a hundred broad, unbranded, disease-awareness questions — the kind every competitor is also trying to own — will naturally return a low one. Neither number is “better.” They are not measuring the same thing.
Why raw numbers mislead leadership
Executives instinctively treat a percentage as an absolute performance signal, a pattern documented long before GEO existed: people anchor on a headline figure and neglect the base rate that produced it2. Applied to GEO, this means a 40% score is read as “four times better” than a 10% score, when in reality the two brand managers may simply have built structurally different — and structurally incomparable — question pools. Two teams optimizing the same product line, using the same content, could report wildly different citation rates purely because one chose easier questions. Comparing them head-to-head rewards the person who picked the softest benchmark, not the person who produced the most citable content.
The fix: calibrate a baseline before you compare
A comparable citation index requires a shared, calibrated instrument, the same principle that underlies any credible cross-brand benchmark in market research. Before two brand managers — or two markets, or two review cycles — compare citation rates, their question pools should include a common baseline sub-pool of broad, high-difficulty, non-branded questions on which no single brand is expected to be cited often.
My practical rule of thumb: if any brand’s citation rate on that baseline pool exceeds roughly 5%, the pool is not difficult or broad enough to serve as a calibration reference, and the comparison should not be trusted.
Once both teams’ baseline sub-pools independently land under that threshold, their headline citation rates — measured on the brand-specific pool built on top of that shared baseline — become meaningfully comparable, because the instrument itself has been standardized. This is analogous to fair-balance discipline in promotional review: a number only means something once its measurement conditions are disclosed and controlled.
Never benchmark against a competitor’s announced number
The same logic rules out comparing your citation rate to a figure a competitor publishes in a press release or case study. Unless that competitor discloses their full question pool, the engines queried, the sampling window, and their baseline calibration, their number is not measuring the same instrument as yours. A competitor citing an eye-catching rate has, most likely, simply built an easier pool — intentionally or not. Treat externally announced GEO figures the way a careful reader treats any statistic without a disclosed methodology: interesting, not actionable, and never a target to chase without first replicating the method2.
What to do before your next steering committee
Before comparing a citation rate across brands, teams, or competitors, agree on three things in writing: the shared baseline sub-pool and its under-5% calibration target, and the engines and sampling window used. Only then does the headline number mean what leadership assumes it means.
References
- Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’24) (pp. 5030–5041). Association for Computing Machinery.
- Tversky A, Kahneman D. Judgment under Uncertainty: Heuristics and Biases. Science. 1974 Sep 27;185(4157):1124-31.
Olivier Gryson, PharmD, MSc
25 years of experience in digital marketing in the pharmaceutical industry
Special focus on AI Search in Pharma Marketing
Frequently Asked Questions
This article was written with the assistance of generative AI technology and reviewed for accuracy.
