Skip to main navigation Skip to search Skip to main content

Artificial Intelligence Chatbots for Dysphagia Patient Education: A Multi-Center International Expert Evaluation

  • Luisa Bertin
  • , Afrin K. Rahman
  • , John Clarke
  • , Amir Mari
  • , Frank Zerbib
  • , C. Prakash Gyawali
  • , Silvia Carrión
  • , Jérôme R. Lechien
  • , Sabine Roman
  • , Elisa Marabotto
  • , Ram Dickman
  • , Vincenzo Savarino
  • , Edoardo Vincenzo Savarino
  • University of Padua
  • Stanford University
  • Nazareth Hospital EMMS
  • Bordeaux University Hospital
  • Washington University St. Louis
  • Centro de Investigación Biomédica en Red
  • Hospital de Mataró
  • Elsan Polyclinique de Poitiers
  • Hospices civils de Lyon
  • University of Genoa
  • San Martino Hospital Genoa
  • Tel Aviv University

Research output: Contribution to journalArticlepeer-review

Abstract

Background: Patients increasingly consult artificial intelligence (AI) chatbots for health information, yet the reliability and accessibility of AI-generated content for complex conditions like dysphagia remain unvalidated. We conducted the first head-to-head comparison of leading large language models for dysphagia patient education, evaluated by an international expert panel. Methods: Forty-six validated questions across four clinical domains were submitted to ChatGPT-4.0 (OpenAI) and Claude 3.7 (Anthropic) in March 2025. Ten blinded experts from six countries rated responses for scientific accuracy (5-point Likert), clarity (5-point Likert), and misinformation (binary). Readability was assessed using Flesch Reading Ease, Flesch–Kincaid Grade Level, and SMOG Index. Between-model comparisons used Wilcoxon signed-rank tests with Cohen's d effect sizes. Key Results: No significant differences emerged for scientific accuracy (ChatGPT: 3.87 ± 0.36 vs. Claude: 3.93 ± 0.35; p = 0.26; d = 0.16), clarity (4.12 ± 0.34 vs. 4.15 ± 0.27; p = 0.67; d = 0.11), or mean misinformation rates (both 2.15; p = 0.96). Strong inter-model correlation existed for accuracy (rs = 0.678; p < 0.001). Critically, both models produced content far exceeding recommended readability levels: SMOG indices of 14.95 ± 2.40 years (ChatGPT) and 17.37 ± 2.67 years (Claude) required extensive education (p < 0.001; d = 0.95) versus the recommended 6–7 years. Categorical analysis showed Claude generated three times more misinformation-free responses (19.6% vs. 6.5%; p = 0.077). Conclusions and Inferences: Leading AI chatbots demonstrate equivalent, acceptable accuracy for dysphagia information but produce content inaccessible to most patients due to excessive complexity. The strong inter-model correlation suggests shared limitations in medical training data. Before clinical implementation, AI-generated patient education requires mandatory readability optimization to address the substantial health literacy gap identified in this study.

Original languageEnglish
Article numbere70344
JournalNeurogastroenterology and Motility
Volume38
Issue number5
DOIs
StatePublished - May 2026

Bibliographical note

Publisher Copyright:
© 2026 John Wiley & Sons Ltd.

Keywords

  • artificial intelligence
  • chatbots
  • deglutition disorders
  • health literacy
  • patient education as topic
  • readability

Fingerprint

Dive into the research topics of 'Artificial Intelligence Chatbots for Dysphagia Patient Education: A Multi-Center International Expert Evaluation'. Together they form a unique fingerprint.

Cite this