Skip to content

Olivier Gryson

  • Home
  • Pharma Marketing in the Age of AI Search
  • Glossary
  • Contact
  • About
Olivier Gryson

The Hidden Bias in Your AI Visibility Score: A Pharma GEO Guide

A brand manager tells the steering committee their product is cited in 40% of relevant AI answers. A colleague, on a different product, reports 10%. Leadership concludes one team has mastered Generative Engine Optimization (GEO) and the other has not. That conclusion is almost certainly wrong.

Citation rate measures the question pool, not the brand

GEO citation rate is not an intrinsic property of a brand or a piece of content. It is the output of a specific instrument: the pool of questions used to query the generative engines, and the sample of engines and prompts run against it. The foundational GEO research from Aggarwal showed that specific content tactics — adding statistics, citing authoritative sources, quoting experts — can shift citation prominence by tens of percentage points depending on the query set tested1. If tactics alone can move the number by 40 points, the composition of the question pool moves it by even more. A pool built from ten narrow, branded, low-competition questions will naturally return a high citation rate. A pool built from a hundred broad, unbranded, disease-awareness questions — the kind every competitor is also trying to own — will naturally return a low one. Neither number is “better.” They are not measuring the same thing.

Why raw numbers mislead leadership

Executives instinctively treat a percentage as an absolute performance signal, a pattern documented long before GEO existed: people anchor on a headline figure and neglect the base rate that produced it2. Applied to GEO, this means a 40% score is read as “four times better” than a 10% score, when in reality the two brand managers may simply have built structurally different — and structurally incomparable — question pools. Two teams optimizing the same product line, using the same content, could report wildly different citation rates purely because one chose easier questions. Comparing them head-to-head rewards the person who picked the softest benchmark, not the person who produced the most citable content.

The fix: calibrate a baseline before you compare

A comparable citation index requires a shared, calibrated instrument, the same principle that underlies any credible cross-brand benchmark in market research. Before two brand managers — or two markets, or two review cycles — compare citation rates, their question pools should include a common baseline sub-pool of broad, high-difficulty, non-branded questions on which no single brand is expected to be cited often.

My practical rule of thumb: if any brand’s citation rate on that baseline pool exceeds roughly 5%, the pool is not difficult or broad enough to serve as a calibration reference, and the comparison should not be trusted.

Once both teams’ baseline sub-pools independently land under that threshold, their headline citation rates — measured on the brand-specific pool built on top of that shared baseline — become meaningfully comparable, because the instrument itself has been standardized. This is analogous to fair-balance discipline in promotional review: a number only means something once its measurement conditions are disclosed and controlled.

Never benchmark against a competitor’s announced number

The same logic rules out comparing your citation rate to a figure a competitor publishes in a press release or case study. Unless that competitor discloses their full question pool, the engines queried, the sampling window, and their baseline calibration, their number is not measuring the same instrument as yours. A competitor citing an eye-catching rate has, most likely, simply built an easier pool — intentionally or not. Treat externally announced GEO figures the way a careful reader treats any statistic without a disclosed methodology: interesting, not actionable, and never a target to chase without first replicating the method2.

What to do before your next steering committee

Before comparing a citation rate across brands, teams, or competitors, agree on three things in writing: the shared baseline sub-pool and its under-5% calibration target, and the engines and sampling window used. Only then does the headline number mean what leadership assumes it means.


References

  1. Aggarwal, P., Murahari, V., Rajpurohit, T., Kalyan, A., Narasimhan, K., & Deshpande, A. (2024). GEO: Generative engine optimization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD ’24) (pp. 5030–5041). Association for Computing Machinery.
  2. Tversky A, Kahneman D. Judgment under Uncertainty: Heuristics and Biases. Science. 1974 Sep 27;185(4157):1124-31. 

Olivier Gryson, PharmD, MSc
25 years of experience in digital marketing in the pharmaceutical industry
Special focus on AI Search in Pharma Marketing


Frequently Asked Questions

Because citation rate is not a property of the brand — it’s a property of the question pool used to measure it. A pool of ten narrow, branded questions naturally returns a higher rate than a pool of a hundred broad, unbranded, disease-awareness questions that every competitor is also trying to own. A 40% score and a 10% score can both be “correct” and still be measuring two different instruments.

Not on its own. A high rate can simply mean the question pool was easier (narrower, more branded, less contested), not that the content or optimization work was superior. Without knowing what pool produced the number, “higher” and “better” aren’t the same claim.

Both pools need a shared, calibrated baseline sub-pool: broad, high-difficulty, non-branded questions where no brand is expected to score high. As my practical rule of thumb, if any brand’s citation rate on that baseline pool exceeds roughly 5%, the pool isn’t difficult enough to serve as a calibration reference, and the comparison shouldn’t be trusted. Once both baselines land under that threshold, the brand-specific rates built on top of them become meaningfully comparable.

No — not unless they disclose their full question pool, the engines and sampling window used, and their baseline calibration. An externally announced GEO figure without a disclosed methodology is not measuring the same instrument as yours, so it isn’t a valid target to chase.

Three things, in writing: the shared baseline sub-pool and its under-5% calibration target, and the engines and sampling window used to generate the numbers. Only once those are fixed does a headline citation rate mean what leadership assumes it means.

Follow the conversation on LinkedIn

I regularly share reflections on pharma marketing, search behavior, and the impact of AI on healthcare communication.

Follow me on LinkedIn

This article was written with the assistance of generative AI technology and reviewed for accuracy.

Related

Published on: August 10, 2026

© 2026 Olivier Gryson - Terms of Use and Privacy - Contact

Content on this website is provided for informational and thought-leadership purposes only. All examples, scenarios, and recommendations are illustrative and intended to stimulate discussion, not to provide medical, legal, regulatory, or compliance advice.

Any pharmaceutical activities must be conducted in accordance with applicable laws and regulations, relevant industry codes of practice (including those of EFPIA and IFPMA), and internal Medical, Legal, and Regulatory (MLR) review and approval processes. Responsibility for compliance remains with the reader and their organization.

  • Home
  • Pharma Marketing in the Age of AI Search
  • Glossary
  • Contact
  • About