Objective: The increasing use of large language models (LLMs) as sources of health information raises concerns regarding the quality and readability of patient-directed content. This study aimed to evaluate and compare the readability and quality of responses generated by three publicly accessible LLMs, ChatGPT-4o, Google Gemini 2.0 Flash, and DeepSeek-R1, to frequently asked patient questions related to periodontology. Materials and Methods: In this cross-sectional study, 48 real-world periodontal questions were retrieved from Reddit and Quora and entered verbatim into each chatbot (February 2025, default settings). Readability was assessed using Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL). Quality and accuracy were evaluated using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool by three independent expert periodontists. Mean scores were compared using one-way ANOVA with Tukey post-hoc tests (α = 0.05). Results: FKGL differed significantly among models (p = 0.007), with ChatGPT-4o producing the most complex text (11.18 ± 1.55) compared to DeepSeek-R1 (9.84 ± 1.78) and Gemini (9.91 ± 1.89). FRE did not differ significantly (p = 0.144), although all scores (42.37–49.55) fell within the “difficult” readability range. Mean QAMAI scores were comparable across models (p = 0.078), ranging from 18.31 ± 1.90 (Gemini) to 19.10 ± 1.78 (DeepSeek-R1), placing the mean scores at the lower end of the predefined “good quality” range. Subgroup analysis revealed significant readability differences only within the “pockets” domain (p = 0.001), while quality remained broadly consistent across topics. Conclusions: All evaluated LLMs produced responses with mean QAMAI scores at the lower end of the predefined good-quality range; however, their readability exceeded recommended standards for public health materials. While LLMs may serve as supplementary educational tools, language simplification strategies and continued professional oversight remain essential to ensure safe and accessible patient information.
Readability and Quality of Chatbot Responses to Periodontal Patient Queries: A Cross-Sectional Evaluation of Three Publicly Accessible Large Language Models
Valente, Nicola Alberto
Primo
;
2026-01-01
Abstract
Objective: The increasing use of large language models (LLMs) as sources of health information raises concerns regarding the quality and readability of patient-directed content. This study aimed to evaluate and compare the readability and quality of responses generated by three publicly accessible LLMs, ChatGPT-4o, Google Gemini 2.0 Flash, and DeepSeek-R1, to frequently asked patient questions related to periodontology. Materials and Methods: In this cross-sectional study, 48 real-world periodontal questions were retrieved from Reddit and Quora and entered verbatim into each chatbot (February 2025, default settings). Readability was assessed using Flesch Reading Ease (FRE) and Flesch–Kincaid Grade Level (FKGL). Quality and accuracy were evaluated using the Quality Analysis of Medical Artificial Intelligence (QAMAI) tool by three independent expert periodontists. Mean scores were compared using one-way ANOVA with Tukey post-hoc tests (α = 0.05). Results: FKGL differed significantly among models (p = 0.007), with ChatGPT-4o producing the most complex text (11.18 ± 1.55) compared to DeepSeek-R1 (9.84 ± 1.78) and Gemini (9.91 ± 1.89). FRE did not differ significantly (p = 0.144), although all scores (42.37–49.55) fell within the “difficult” readability range. Mean QAMAI scores were comparable across models (p = 0.078), ranging from 18.31 ± 1.90 (Gemini) to 19.10 ± 1.78 (DeepSeek-R1), placing the mean scores at the lower end of the predefined “good quality” range. Subgroup analysis revealed significant readability differences only within the “pockets” domain (p = 0.001), while quality remained broadly consistent across topics. Conclusions: All evaluated LLMs produced responses with mean QAMAI scores at the lower end of the predefined good-quality range; however, their readability exceeded recommended standards for public health materials. While LLMs may serve as supplementary educational tools, language simplification strategies and continued professional oversight remain essential to ensure safe and accessible patient information.| File | Dimensione | Formato | |
|---|---|---|---|
|
International Journal of Dentistry - 2026 - Valente - Readability and Quality of Chatbot Responses to Periodontal Patient.pdf
accesso aperto
Descrizione: Articolo principale
Tipologia:
versione editoriale (VoR)
Dimensione
520.19 kB
Formato
Adobe PDF
|
520.19 kB | Adobe PDF | Visualizza/Apri |
I metadati presenti in IRIS UNICA sono rilasciati con licenza Creative Commons CC0 1.0 Universal, mentre i file delle pubblicazioni sono protetti da diritto d'autore, salvo diversa indicazione.



