Skip to content

Olivier Gryson

  • Home
  • Pharma Marketing in the Age of AI Search
  • Glossary
  • Contact
  • About
Olivier Gryson

GEO for Rare Diseases: Unique Opportunities and Challenges

For rare diseases, visibility is no longer only about ranking on Google — it is about becoming a trusted source that AI assistants cite.

Why Is GEO Particularly Important for Rare Diseases?

The Rarity Paradox

Rare diseases are anything but rare at the population level. According to The Lancet Global Health, more than 7,000 types of rare disease exist worldwide, collectively affecting an estimated 300 million people1 — a figure that, on a global population of approximately 8 billion, represents close to 4% of humanity. Despite this collective burden, each individual condition may affect only a few thousand patients, creating a structural paradox: the diseases are common enough to represent a major public health problem, yet too uncommon for traditional search engines to surface meaningful, consistent information about them2.

Traditional Search Engine Optimization (SEO) is built on search volume: the more people query a term, the more competitive and well-indexed it becomes. For rare diseases, search volume is structurally low. A condition affecting 1 in 100,000 individuals generates a fraction of the queries that common conditions generate. This means standard SEO ranking algorithms deprioritize rare disease content in comparison to high-traffic health topics. As a consequence, when a patient or caregiver searches for answers about a rare condition, they frequently encounter fragmented, outdated, or scientifically unreliable results. The information ecosystem for rare diseases is underdeveloped not because the need is absent, but because the economic incentives of search-volume-driven SEO do not reward content investment in this domain.

The Long Diagnostic Journey

The challenge of information visibility has direct clinical consequences. A landmark retrospective patient survey published in the European Journal of Human Genetics — the EURORDIS Rare Barometer study by Faye et al. (2024), conducted among 6,507 patients across 41 countries and encompassing 1,675 distinct rare diseases — found that the average total time from symptom onset to confirmed diagnosis is 4.7 years, and that 56% of respondents waited more than six months after first medical contact before receiving a confirmed diagnosis3. A separate, larger EURORDIS Rare Barometer survey spanning more than 10,000 patients from 42 countries found that 25% of respondents needed eight or more consultations with a healthcare professional before receiving a definitive diagnosis, and that approximately 60% were initially misdiagnosed — with many having their symptoms dismissed as psychological or attributed to unrelated physical conditions4.

Among children and adolescents the situation is even more acute. Faye et al. analysis3 found that symptom onset before age 30 was the strongest determinant of diagnostic delay, with children facing an odds ratio of 3.11 and adolescents an odds ratio of 4.79 for prolonged diagnostic journeys — both relative to adult-onset patients in the same dataset. Separate US-based data from an undiagnosed rare disease clinic reported a median time from symptom onset to specialist referral of 7.6 years5. Studies of specific conditions illustrate the extremes: a retrospective analysis of patients with Ehlers-Danlos Syndrome found that those who received an initial misdiagnosis waited an average of 19 years for a correct diagnosis, compared to 8 years for those who were not initially misdiagnosed6.

AI Search Changes Discovery

The rise of generative AI is fundamentally restructuring how patients, caregivers, and clinicians find medical information. According to figures disclosed by OpenAI in January 2026, more than 40 million people submit healthcare-related questions to ChatGPT every day, healthcare topics account for more than 5% of all messages sent to the platform globally, and approximately 70% of health-related interactions occur outside typical clinic hours — underscoring that AI assistants are increasingly becoming the first point of contact when professional guidance is unavailable7. These figures reflect OpenAI’s own reporting and have not been independently validated in peer-reviewed literature.

For physicians, AI adoption is accelerating rapidly. According to the American Medical Association (AMA), 66% of US physicians reported using AI for at least one work-related use case in 2024, up from 38% in 2023 — a year-over-year increase drawn from comparable waves of the same AMA survey. Among patients who reported using AI tools to manage their health, 55% used them to check or explore symptoms, 48% used them to understand medical terms or instructions, and 44% used them to understand treatment options8.

This behavioral shift is particularly consequential for rare diseases. When a patient types a conversational description of unusual symptoms into ChatGPT, the AI does not return a ranked list of links — it synthesizes an answer drawn from the sources it considers most authoritative and well-structured. If a rare disease organization or pharmaceutical company has invested in entity-rich, well-structured content that is legible to AI models, that content may surface prominently. If it has not, the patient may receive incomplete information, or none at all.

Why LLMs Can Bridge Fragmented Information

Large language models have a structural advantage over keyword-based search in the rare disease context: they can reason across heterogeneous sources, integrate disparate terminologies, and synthesize information from clinical literature, patient advocacy resources, and medical databases. For diseases with evolving nomenclature, multiple synonyms, and scattered scientific literature, LLMs can theoretically draw connections that a traditional search engine would miss. A patient describing ‘muscle weakness and breathing problems in a baby’ may not know to search for Spinal Muscular Atrophy or Pompe disease; an LLM may be able to bridge that gap.

However, this bridging capacity is entirely dependent on the quality, structure, and authoritative weight of the content that LLMs are trained on and retrieve from. Generative Engine Optimization (GEO) — defined by Aggarwal et al. as the practice of structuring and optimizing content so that AI systems can find it, understand it, and cite it in their responses — is therefore not a peripheral concern for rare disease stakeholders. It is central to patient discovery, disease education, and ultimately to shortening the diagnostic odyssey.9

Key Takeaway — For rare diseases, visibility is no longer only about ranking on Google. It is about becoming a trusted source that AI assistants cite, summarize, and return to patients and clinicians who are searching for answers they cannot find elsewhere.

What Makes Rare Disease Content Difficult for AI Models to Understand?

While LLMs offer genuine potential for rare disease information synthesis, systematic research reveals significant limitations in their current ability to reason reliably about these conditions. Understanding these weaknesses is a prerequisite for building content that compensates for them.

A Sparse and Contested Knowledge Ecosystem

The fundamental problem is one of data volume. LLMs learn from the corpus of content available on the internet and in scientific literature. For common diseases such as diabetes or hypertension, this corpus is vast: thousands of peer-reviewed articles, clinical guidelines, patient-facing explanations, and educational materials. For rare diseases, particularly those affecting fewer than one in 100,000 individuals, the available corpus may consist of a handful of case reports, a single clinical guideline, and a few patient advocacy pages. When the knowledge ecosystem is weak, LLMs produce outputs that are correspondingly uncertain, inconsistent, or frankly inaccurate.

A systematic assessment of six leading LLMs — including GPT-4, Claude 3.7, Llama-3.3 70B, and Gemma-2 27B — evaluated their ability to generate core phenotypic features and causal genes for 10,892 Orphanet diseases. The study found that LLM recall of curated rare disease knowledge was generally low. While commercial models achieved over 60% recall for gene associations, they struggled considerably with precise phenotype recovery. The authors concluded that current generalist LLMs lack the precision to replace curated rare disease knowledge bases, but offer complementary, semantically relevant information, and recommended hybrid approaches combining expert curation with selectively integrated LLM data10.

Conflicting Terminology, Synonyms, and Evolving Nomenclature

Rare diseases are afflicted with a particularly severe version of a problem that undermines AI understanding: nomenclature fragmentation. Consider Pompe disease. Its full list of recognized synonyms includes Glycogen Storage Disease Type II (GSD II), Glycogen Storage Disease Due to Acid Maltase Deficiency, Acid Maltase Deficiency, Acid Alpha-Glucosidase Deficiency, GAA Deficiency, Glycogenosis Type II, and Alpha-1,4-glucosidase acid deficiency. It holds the Orphanet code ORPHA:365 and is classified under ICD-10 code E74.0. Depending on the source — whether a research paper from 1985, a current clinical guideline, a patient forum, or an orphan drug application — any of these terms may appear without cross-referencing the others.11, 12

The same issue applies to Amyotrophic Lateral Sclerosis (ALS), also known as Lou Gehrig’s disease and, throughout the European literature, as Motor Neurone Disease (MND). A model trained predominantly on US content may associate ALS primarily with its eponym and acronym, while failing to draw equivalent inferential connections when a British or Australian user refers to MND. A study applying LLMs to rare disease differential diagnosis at scale found only moderate performance — approximately 56.8% Top-5 accuracy across a broad set of rare conditions — partly attributable to this terminological fragmentation, and representing a performance baseline that would require substantial improvement before clinical application.13

Sparse Entity Relationships and Limited Authoritative Sources

AI models do not simply retrieve documents; they build internal representations of entities and the relationships between them. For a condition like Type 2 Diabetes, the semantic graph is dense: there are well-established relationships between the disease entity, its biomarkers (HbA1c), its diagnostic criteria (WHO 2006 criteria), its first-line treatments (metformin), its complications (diabetic nephropathy), its genetic susceptibility loci, and its epidemiological profile. For most rare diseases, this semantic graph is sparse.

A study developing a Facial Phenotype Knowledge Graph for rare genetic diseases found that even after constructing a purpose-built graph with 6,143 nodes and 19,282 relationships, LLMs remained prone to hallucination and limited domain knowledge when applied to rare genetic disease queries. Notably, Retrieval-Augmented Generation (RAG) reduced temperature-induced variability in model responses by 53.94% within that specific experimental framework — a finding that suggests RAG meaningfully improves output consistency, though results should be interpreted in the context of the study’s specific disease set and model configurations.14

For pharmaceutical companies and medical organizations active in the rare disease space, the practical implication is direct: if your content does not make explicit the semantic connections between a disease, its synonyms, its genetic basis, its symptoms in lay language, its diagnostic pathway, and its treatment options, AI models will fail to retrieve and synthesize it accurately — regardless of its scientific quality.

Key Takeaway — LLMs struggle with rare disease content not because they are unintelligent, but because the knowledge ecosystem they depend on is fragmented, terminologically inconsistent, and sparsely interconnected. Addressing this requires deliberate content architecture, not simply more content.

How Can Pharmaceutical Companies Create AI-Citable Rare Disease Content?

This section constitutes the practical core of a rare disease GEO strategy. The goal is not to produce more content. It is to produce the right kind of content — structured to be understood by AI models, rich in medical entities, and comprehensive enough to serve as a primary reference.

Build Entity-Rich Content That Maps the Full Disease Landscape

Rather than creating isolated pages about a single treatment or a clinical trial, pharmaceutical companies should build content that positions the disease as its primary subject. An AI model retrieving content about Fabry disease, for example, should find a resource that covers the following interconnected entities:15,16

  • Disease definition and classification: What Fabry disease is, its Orphanet code (ORPHA:324 [15]), its ICD-10 classification (E75.21), its status as an X-linked lysosomal storage disorder.
  • Genetics: Mutations in the GLA gene encoding alpha-galactosidase A, X-linked inheritance pattern, penetrance differences in hemizygous males versus heterozygous females, known pathogenic variants.
  • Symptoms: Described both in clinical terminology (angiokeratoma, acroparesthesia, hypohidrosis, cardiomyopathy, nephropathy) and in plain patient language (‘burning pain in hands and feet’, ‘dark red skin spots’, ‘kidney problems’, ‘heart thickening’).
  • Epidemiology: Estimated prevalence of classical Fabry disease of approximately 1 in 40,000 to 1 in 60,000 in the general population; newborn screening studies have identified higher rates of later-onset variants in some populations [16]. Geographic variation and founder mutation effects should be noted where relevant.
  • Diagnosis: How the disease is confirmed — enzyme activity assay (alpha-Gal A), genetic testing, biomarkers including lyso-Gb3, and the conditions under which newborn screening identifies affected individuals.
  • Patient pathway: The typical journey from first symptom to referral to a metabolic specialist, common misdiagnoses (multiple sclerosis, growing pains, anxiety disorder), average diagnostic delay.
  • Treatment options: Approved therapies (enzyme replacement therapy with agalsidase alfa or agalsidase beta; pharmacological chaperone migalastat for eligible patients with amenable GLA variants), mechanism of action, administration route, and evidence base.
  • Support organizations: National and international patient advocacy organizations, centers of expertise, registries.

This entity architecture does two things. First, it creates a document that functions as a knowledge hub rather than a marketing page, which is what AI retrieval systems reward. Second, it explicitly mirrors the semantic graph that an LLM needs in order to answer diverse user questions about the condition — from a patient asking about symptoms to a physician seeking diagnostic criteria to a genetic counselor asking about carrier status in women.

Answer Complete Questions, Not Just Topics

One of the most consistent findings from GEO research is that AI systems prefer content that directly answers full questions over content that lists topics. This is because users of generative AI ask questions in natural language, and the AI retrieves content that matches the structure of those questions. The practical implication for rare disease content is a shift from topic headers to question headers.

Instead of a section titled ‘Symptoms’, structure content around:

  • What are the first symptoms of Fabry disease in children?
  • What are the early warning signs of Pompe disease in adults?
  • What symptoms distinguish ALS from other motor neuron diseases?

Instead of ‘Diagnosis’, structure content around:

  • How is Fabry disease diagnosed in adults?
  • What tests confirm a diagnosis of Pompe disease?
  • What biomarkers are used to diagnose lysosomal storage disorders?

This approach serves both GEO and traditional SEO simultaneously, because natural language queries — typed or spoken — now account for an increasing share of health searches. Each question-structured section can stand alone as an answer unit, making it easier for AI models to retrieve specific components without needing to process an entire long-form document.

Explain Concepts in Multiple Languages — Scientific, Clinical, and Plain

Rare disease content must serve multiple audiences simultaneously: patients who know nothing about their condition at time of first search, caregivers with some medical literacy, primary care physicians who need clinical orientation, and specialists who need technical precision. Because AI models serve all of these audiences, content that exists in only one register will systematically fail some users.

The optimal content architecture uses layered explanations: a plain-language summary of each key concept followed by clinical elaboration and scientific detail. For example, a section on Pompe disease muscle involvement might be structured as follows:

Plain language:In Pompe disease, muscles throughout the body become progressively weaker because the body cannot properly break down a sugar called glycogen. This affects the muscles used for breathing, walking, and posture.
Clinical language:Pompe disease (GSD II) is characterized by progressive myopathy resulting from lysosomal glycogen accumulation due to GAA enzyme deficiency. Respiratory muscle involvement leads to diaphragmatic weakness and restrictive lung disease, while skeletal muscle pathology produces proximal limb-girdle weakness.
Scientific language:Deficiency of acid alpha-glucosidase (GAA; EC 3.2.1.20) causes lysosomal glycogen accumulation in type 1 and type 2 muscle fibers, resulting in autophagosome dysfunction, disrupted calcium homeostasis, and progressive sarcomere disruption.

This layered approach is not redundant — each layer addresses a different query pattern. The plain language layer matches conversational AI queries from patients and families. The clinical layer matches queries from primary care physicians. The scientific layer matches queries from specialists and researchers.


How Can GEO Shorten the Diagnostic Journey for Patients?

This section addresses what is arguably the most consequential — and most underappreciated — dimension of rare disease GEO: its potential to contribute to improved patient outcomes by shortening the diagnostic odyssey. Unlike marketing-focused content strategies, a well-executed GEO strategy for rare diseases can serve as a meaningful patient information intervention.

Patients Ask ChatGPT Before They See a Doctor

The behavioral reality of modern healthcare-seeking is that generative AI has become the first point of contact for millions of people describing symptoms they cannot explain. A 2024 Deloitte survey found that approximately 48% of US consumer respondents reported using generative AI for health-related concerns.17 These numbers are even higher in emerging markets. It reaches 85% in India.18

For rare disease patients, this behavioral shift has distinctive implications. Research on rare disease information-seeking demonstrates that patients contacting genetic and rare disease information centers most often seek information about disease prognosis, finding a specialist, and obtaining a diagnosis for symptoms19. A survey of Italian parents of children with rare diseases found that 87% accessed the internet daily, with 99% searching for disease characteristics and 89% searching for diagnostic information20. In the rare disease context, the online search is not supplemental to the diagnostic process — it is often intertwined with it.

Parents Describe Symptoms Conversationally

When a parent observes that their infant has difficulty swallowing, seems unusually floppy, tires quickly while feeding, and has a noticeably large heart on echocardiogram, they are unlikely to search for ‘infantile-onset Pompe disease.’ They will describe what they see, in the language they have. A parent might type: ‘Why is my baby too weak to hold up his head and having trouble breathing?’ or ‘My newborn has a big heart and weak muscles — what disease could cause this?’

If the content ecosystem around infantile-onset Pompe disease has been structured with GEO in mind — explicitly connecting the disease entity to its symptom presentations in plain language, answering conversational questions about early warning signs in infants, and providing diagnostic pathways in accessible terms — then a well-trained AI model should be able to surface this information in response to those conversational queries. If the only available content uses dense clinical terminology without plain-language bridging, the AI will fail this family.

HCPs Exploring Differential Diagnoses

The diagnostic journey does not depend only on patients. Physicians — particularly primary care physicians who serve as gatekeepers to rare disease specialist referral — are increasingly turning to AI tools for clinical decision support. According to the AMA, 66% of US physicians used AI for at least one work-related purpose in 2024.8 Research on LLM use in differential diagnosis for rare diseases, while finding current accuracy limitations, identifies a clear emerging role: LLMs can serve as a first-pass differential diagnosis generation tool that flags rare diseases for specialist consideration.13

A primary care physician seeing a patient with progressive proximal muscle weakness and recurrent respiratory infections, uncertain whether this represents a neuromuscular condition, a metabolic disorder, or an inflammatory process, may query an AI assistant for a differential diagnosis. If content covering late-onset Pompe disease (LOPD) includes explicit clinical features, distinguishing characteristics from similar conditions (myasthenia gravis, Becker muscular dystrophy, idiopathic inflammatory myopathies), recommended diagnostic tests, and referral criteria, that content has a meaningfully higher probability of appearing in the AI’s synthesized response.

AI Summarizing Scattered Evidence

One of the most direct contributions a well-structured rare disease knowledge hub can make is to help AI models produce accurate summaries of scattered evidence. The literature on rare diseases is typically distributed across multiple journals, registries, clinical case series, and orphan drug assessment reports. No single document summarizes the evidence base for most conditions. AI models performing retrieval-augmented generation on behalf of a clinician or patient will aggregate across these sources — but the quality of that aggregation is highly dependent on the structure and authority of the underlying content.

Research on hybrid frameworks combining Orphanet’s Rare Disease Ontology (ORDO) and the Unified Medical Language System (UMLS) with LLMs found that integrating structured ontological frameworks significantly improved rare disease identification from unstructured clinical notes.21 For pharmaceutical companies and medical organizations, this means that content which explicitly uses Orphanet codes, HPO (Human Phenotype Ontology) terms, and UMLS-aligned terminology is not only more academically rigorous — it is more likely to be retrieved, understood, and cited by AI systems.

From Discoverability to Diagnosis: The GEO Patient Impact Chain

The patient impact pathway from effective rare disease GEO works as follows: a patient or caregiver describes symptoms in natural language to an AI assistant; the AI retrieves and synthesizes content from authoritative, well-structured sources; the response includes the name of a rare condition, a description of its diagnostic pathway, and a recommendation to consult a specialist; the patient or caregiver brings this information to a physician; the physician is thereby alerted to a diagnostic possibility they might not have considered; testing is initiated earlier; and diagnosis is reached sooner.

At each step of this chain, the quality of the underlying content is the limiting factor. GEO investment in rare diseases is therefore not only a marketing strategy — it may contribute to the information infrastructure that supports earlier diagnosis. The magnitude of this potential impact has not been formally quantified, and content quality is only one of many factors that determine diagnostic delay.

Key Takeaway — GEO in the rare disease context is not only a communications exercise — it may function as a patient information intervention. Every authoritative, AI-legible content asset has the potential to contribute to earlier diagnosis by surfacing the right information at the right moment in a patient’s journey.

What Are the Best Practices for Building a Rare Disease GEO Strategy?

The following framework translates the insights from Sections 1 through 4 into concrete, actionable recommendations for pharmaceutical companies, patient advocacy organizations, and medical communications teams working in the rare disease space.

1. Build Trusted Knowledge Hubs, Not Isolated Product Pages

The single most consequential structural shift an organization can make is to move away from product-centric content architecture toward disease-centric knowledge hubs. A knowledge hub is a comprehensive, authoritative resource that positions the disease as the central subject, covering all relevant entities and their relationships, rather than using the disease as a vehicle to introduce a specific therapy.

Industry commentary from pharmaceutical marketing networks suggests that AI systems cross-reference multiple authoritative sources before surfacing content, and that brands publishing isolated marketing copy without external validation face challenges in gaining AI visibility. A knowledge hub, by contrast, demonstrates epistemic completeness — the quality that AI systems tend to reward because it allows a generative model to retrieve a single source capable of answering a wide variety of related queries.21

2. Connect Every Medical Entity Through Semantic Relationships

Content architecture for AI legibility requires explicit semantic connections between:

  • Disease name and all recognized synonyms (including historical, regional, and colloquial variants)
  • ORPHA codes, ICD-10/ICD-11 codes, OMIM identifiers, and HPO term sets
  • Causative genes and their protein products
  • Molecular mechanism and resulting pathophysiology
  • Clinical symptoms in both medical and plain-language registers
  • Diagnostic biomarkers and confirmatory tests
  • Approved treatments and their mechanisms
  • Patient organizations, registries, centers of expertise, and clinical trial networks

Research on LLM-based rare disease phenotyping demonstrated that incorporating Orphanet’s Rare Disease Ontology and UMLS vocabulary into content retrieval frameworks substantially improved model accuracy.21 Making these ontological relationships explicit in published content — through structured data markup (schema.org), internal linking, and explicit cross-referencing — extends this benefit to any AI model retrieving your content.

3. Prioritize Quality and Authority Over Quantity

Rare diseases do not need thousands of content assets. They need a small number of exceptional ones. The research literature consistently demonstrates that LLMs reward evidence density, source authority, and citation quality. A single, rigorously referenced, editorially reviewed knowledge hub for a given rare disease — one that explicitly cites clinical guidelines, patient registry data, and orphan drug assessment reports — will outperform dozens of thin, uncited pages.

Structured data implementation is foundational to AI legibility. Research on authority signals in AI-cited health sources found that authorship credentials, institutional affiliations, and editorial review processes are among the strongest predictors of whether a source is cited by AI systems.23 For rare disease content, this means publishing under the byline of credentialed medical authors, citing primary literature, disclosing the date of last review, and using schema.org MedicalCondition and Drug markup to make authority signals machine-readable.

4. Keep Scientific Information Current

LLMs increasingly incorporate recency signals into their retrieval and ranking behavior. For rare diseases, the clinical evidence base is rapidly evolving: new gene therapy approvals, updated clinical guidelines, revised diagnostic criteria, and new patient registry data emerge continuously. Content that was accurate in 2020 may be misleading in 2026. A GEO strategy for rare diseases must include a scheduled content review process — ideally aligned with the publication cycle of relevant clinical guidelines and Orphanet’s periodic ontology updates — to ensure that scientific information reflects the current state of evidence.

Displaying explicit ‘last reviewed’ dates, linking to current clinical guidelines, and flagging when information reflects interim or evolving evidence are not merely best practices for scientific accuracy. They are GEO-relevant signals that communicate content trustworthiness to AI retrieval systems.

5. Measure AI Visibility, Not Only Traditional Search Metrics

The metrics appropriate to a rare disease GEO strategy differ materially from traditional SEO metrics. Given the low search volumes inherent to rare disease queries, organic traffic volumes will never be the primary indicator of content performance. Instead, organizations should monitor24:

  • AI citation rate: How frequently does your content appear as a cited source in responses from ChatGPT, Perplexity, Google AI Overviews, and similar systems?
  • Brand and entity mention rate: How often are your organization’s name, your disease knowledge hub’s URL, and the specific entities you have defined mentioned in AI-generated responses?
  • Entity recognition quality: When an AI model is asked about the disease you have covered, does it correctly identify all synonyms, distinguish your disease from related conditions, and accurately describe the diagnostic pathway?
  • Retrieval quality: When you submit a test query to an AI system, how complete and accurate is the synthesized response, and which specific content elements are reflected in it?
Key Takeaway — A rare disease GEO strategy is a long-term infrastructure investment, not a one-time content project. It requires disease-centric architecture, semantic entity connections, scientific authority, continuous updating, and AI-specific measurement frameworks.

Rare Disease GEO Checklist

Use this checklist when reviewing any rare disease content asset intended for AI discoverability:

  • Is the disease clearly and unambiguously defined, with its Orphanet code, ICD-10/ICD-11 classification, and OMIM identifier stated?
  • Are all recognized synonyms, acronyms, historical names, and regional terminological variants explicitly included?
  • Are symptoms explained in both medical/clinical terminology and plain patient language, covering both early and advanced presentations?
  • Is the diagnostic pathway described end-to-end, from first symptom presentation through confirmatory testing, including biomarkers, imaging criteria, and genetic testing?
  • Are authoritative sources cited throughout, including peer-reviewed publications, clinical guidelines, orphan drug assessment reports, and patient registry data?
  • Are entities linked semantically — including connections between the disease, its causative gene(s), relevant biomarkers, available treatments, and patient organizations?
  • Is structured data markup (schema.org MedicalCondition, Drug, or equivalent) implemented to make authority signals and entity relationships machine-readable?
  • Is the content structured around complete natural-language questions (e.g., ‘How is [disease] diagnosed?’) rather than bare topic headers?
  • Is content layered for multiple audiences — plain language for patients, clinical language for physicians, scientific language for specialists?
  • Does the content include a clearly visible ‘last reviewed’ date and commitment to updating as clinical evidence evolves?
  • Is AI citation and entity mention rate being tracked alongside traditional search metrics?
  • Can each major section answer a specific user question independently, without requiring the reader to consume the full document?

References

  1. The Lancet Global Health. The landscape for rare diseases in 2024. Lancet Glob Health. 2024 Mar;12(3):e341.
  2. Delivering Hope for Rare Diseases https://ncats.nih.gov/sites/default/files/NCATS_RareDiseasesFactSheet.pdf Last accessed 29/06/2026
  3. Faye F, Crocione C, Anido de Peña R, Bellagambi S, Escati Peñaloza L, Hunter A, Jensen L, Oosterwijk C, Schoeters E, de Vicente D, Faivre L, Wilbur M, Le Cam Y, Dubief J. Time to diagnosis and determinants of diagnostic delays of people living with a rare disease: results of a Rare Barometer retrospective patient survey. Eur J Hum Genet. 2024 Sep;32(9):1116-1126.
  4. Major survey reveals lengthy diagnostic delays for rare disease patients, https://www.eurordis.org/survey-reveals-lengthy-diagnostic-delays/ Last accessed 29/06/2026
  5. Proceedings of IMPRS. Diagnostic Delays and Socio-Demographic Factors Among Patients Referred to an Undiagnosed Rare Disease Clinic. doi:10.18060/29689 https://journals.indianapolis.iu.edu/index.php/IMPRS/article/view/29689 Last accessed 29/06/2026
  6. Kvancz DA. Rare Disease Diagnosis and Management. Am J Pharm Benefits. 2016;8(4):128-132. https://www.pharmacytimes.com/view/the-impact-of-rare-diseases-and-drug-therapy Last accessed 29/06/2026
  7. OpenAI. Healthcare on ChatGPT: usage statistics disclosed January 2026, as reported by Healthcare Finance News. healthcarefinancenews.com. January 2026. https://cdn.openai.com/pdf/2cb29276-68cd-4ec6-a5f4-c01c5e7a36e9/OpenAI-AI-as-a-Healthcare-Ally-Jan-2026.pdf Last accessed 29/06/2026
  8. American Medical Association. AMA Digital Medicine Report: Physician AI Adoption Survey 2024. ama-assn.org. As reported by Fierce Healthcare, January 2026. https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf Last accessed 29/06/2026
  9. Aggarwal P, et al. GEO: Generative Engine Optimization. arXiv:2311.09735. Princeton University / University of Maryland. 2023. https://arxiv.org/pdf/2311.09735 Last accessed 29/06/2026
  10. Groza T, Marcello AJ, Carlisle T, Lim WK, Haendel M, Karnani N, Robinson PN, Graessner H, Chong JX, Baynam G, Jamuar SS. A systematic assessment of large language models’ knowledge of rare diseases: How much do large language models know about rare disease? HGG Adv. 2026 Jan 15;7(1):100558.
  11. Sperry E, Leslie N, Berry L, et al. Pompe Disease. 2007 Aug 31 [Updated 2025 Aug 21]. In: Adam MP, Bick S, Mirzaa GM, et al., editors. GeneReviews® [Internet]. Seattle (WA): University of Washington, Seattle; 1993-2026. Available from: https://www.ncbi.nlm.nih.gov/books/NBK1261/
  12. https://www.orpha.net/en/disease/detail/365 Last accessed 29/06/2026
  13. Schumacher E, Naik D, Kannan A. Rare Disease Differential Diagnosis with Large Language Models at Scale: From Abdominal Actinomycosis to Wilson’s Disease. arXiv:2502.15069. 2025. https://proceedings.mlr.press/v298/schumacher25a.html
  14. Song J, Xu Z, He M, Feng J, Shen B. Graph retrieval augmented large language models for facial phenotype associated rare genetic disease. NPJ Digit Med. 2025 Aug 24;8(1):543. doi: 10.1038/s41746-025-01955-x. Erratum in: NPJ Digit Med. 2025 Oct 8;8(1):604.
  15. https://www.orpha.net/en/disease/detail/324 Last accessed 29/06/2026
  16. Germain DP. Fabry disease. Orphanet J Rare Dis. 2010 Nov 22;5:30.
  17. Deloitte. 2024 Global Health Care Consumer Survey. deloitte.com. 2024. https://www.deloitte.com/be/en/Industries/life-sciences-health-care/analysis/global-health-care-outlook.html Last accessed 29/06/2026
  18. Consumers Are Ready for AI-Enabled Health Care. Health Systems Need to Be, Too. https://www.bcg.com/publications/2026/consumers-are-ready-for-ai-health-care-are-systems Last accessed 29/06/2026
  19. Morgan T, Schmidt J, Haakonsen C, Lewis J, Della Rocca M, Morrison S, Biesecker B, Kaphingst KA. Using the internet to seek information about genetic and rare diseases: a case study comparing data from 2006 and 2011. JMIR Res Protoc. 2014 Feb 24;3(1):e10. 
  20. Tozzi AE, Mingarelli R, Agricola E, Gonfiantini M, Pandolfi E, Carloni E, Gesualdo F, Dallapiccola B. The internet user profile of Italian families of patients with rare diseases: a web survey. Orphanet J Rare Dis. 2013 May 16;8:76. 
  21. Wu J, Dong H, Li Z, Wang H, Li R, Patra A, Dai C, Ali W, Scordis P, Wu H. A hybrid framework with large language models for rare disease phenotyping. BMC Med Inform Decis Mak. 2024 Oct 8;24(1):289.
  22. Pharma Marketing Network. AI Content Optimization Pharma Strategies for 2026. pharma-mkting.com. May 2026. [Industry commentary; not peer-reviewed]
  23. Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses. arXiv:2601.17109. 2025. https://arxiv.org/pdf/2601.17109
  24. https://searchengineland.com/what-is-generative-engine-optimization-geo-444418 Last accessed 29/06/2026

Olivier Gryson, PharmD, MSc
25 years of experience in digital marketing in the pharmaceutical industry
Special focus on AI Search in Pharma Marketing


Frequently Asked Questions

Despite each condition being individually uncommon, rare diseases collectively affect an estimated 300 million people globally — roughly 4% of humanity — across more than 7,000 distinct conditions (The Lancet Global Health, 2024).1

On average, 4.7 years from symptom onset to confirmed diagnosis. Around 56% of patients wait more than six months after first medical contact, and approximately 60% are initially misdiagnosed (Faye et al., European Journal of Human Genetics, 2024).3

Yes. Children face an odds ratio of 3.11 and adolescents 4.79 for prolonged diagnostic journeys compared to adult-onset patients. In some cases, such as Ehlers-Danlos Syndrome, misdiagnosed patients waited an average of 19 years for a correct diagnosis (Faye et al., 2024; Kvancz, 2016).3

Over 40 million people submit healthcare-related questions to ChatGPT daily, with approximately 70% of those interactions occurring outside clinic hours. Among patients using AI for health, 55% use it to explore symptoms and 44% to understand treatment options (OpenAI, January 2026).7

Performance remains limited. A systematic assessment of six leading LLMs — including GPT-4 and Claude 3.7 — found generally low recall of rare disease knowledge, with commercial models achieving over 60% recall for gene associations but struggling considerably with precise phenotype recovery (Groza et al., HGG Advances, 2026).10

Rare diseases frequently carry multiple names across different eras and regions. Pompe disease alone has at least seven recognized synonyms. A study on LLM-based rare disease differential diagnosis found only approximately 56.8% Top-5 accuracy, partly attributable to this terminological fragmentation (Schumacher et al., 2025).13

Evidence suggests it can help. In one study using a purpose-built Facial Phenotype Knowledge Graph, RAG reduced temperature-induced variability in model responses by 53.94%, meaningfully improving output consistency — though results should be interpreted in the context of that specific study’s design (Song et al., NPJ Digital Medicine, 2025).14

66% of US physicians reported using AI for at least one work-related purpose in 2024, up from 38% in 2023 — a near-doubling in a single year — according to the American Medical Association’s physician AI sentiment survey.8

Research consistently shows high online health-seeking behaviour in this population. A survey of Italian families of children with rare diseases found 87% accessed the internet daily, with 99% searching for disease characteristics and 89% seeking diagnostic information (Tozzi et al., Orphanet Journal of Rare Diseases, 2013).20

Yes. Research on hybrid frameworks combining Orphanet’s Rare Disease Ontology (ORDO) and the Unified Medical Language System (UMLS) with LLMs found that integrating structured ontological frameworks significantly improved rare disease identification from unstructured clinical notes (Wu et al., BMC Medical Informatics and Decision Making, 2024).21

Follow the conversation on LinkedIn

I regularly share reflections on pharma marketing, search behavior, and the impact of AI on healthcare communication.

Follow me on LinkedIn

This article was written with the assistance of generative AI technology and reviewed for accuracy.

Related

Published on: June 29, 2026

© 2026 Olivier Gryson - Terms of Use and Privacy - Contact

Content on this website is provided for informational and thought-leadership purposes only. All examples, scenarios, and recommendations are illustrative and intended to stimulate discussion, not to provide medical, legal, regulatory, or compliance advice.

Any pharmaceutical activities must be conducted in accordance with applicable laws and regulations, relevant industry codes of practice (including those of EFPIA and IFPMA), and internal Medical, Legal, and Regulatory (MLR) review and approval processes. Responsibility for compliance remains with the reader and their organization.

  • Home
  • Pharma Marketing in the Age of AI Search
  • Glossary
  • Contact
  • About