Skip to main navigation Skip to search Skip to main content

Conducting Medication Reviews: A Comparative Study Between ChatGPT-4 and Healthcare Professionals

  • Simone M.K. ten Hoope
  • , Sabrina Marongiu
  • , Carl E.H. Siegert
  • , Eibert R. Heerdink
  • , Marjo J.A. Janssen
  • , Fatma Karapinar-Çarkit*
  • *Corresponding author for this work
  • Onze Lieve Vrouwe Gasthuis

Research output: Contribution to journalArticleAcademicpeer-review

Abstract

Background: The increasing prevalence of patients with hyperpolypharmacy (> 10 medications) has made medication reviews increasingly complex. ChatGPT-4-Turbo (ChatGPT), a large language model, has demonstrated potential in healthcare applications and could potentially support medication reviews. Objectives: This study aimed to evaluate the agreement between medication reviews conducted by ChatGPT compared to healthcare professionals (HCPs) in older people. Secondary objectives included: the validity of additional interventions detected by ChatGPT, its ability to structure diagnoses to medication use, and laboratory target values based on patient characteristics. Methods: In this retrospective proof-of-concept study, 51 medication reviews previously conducted by a geriatric internist and hospital pharmacist were re-evaluated using ChatGPT. ChatGPT was trained on polypharmacy guidelines, was then provided with the same primary data as HCPs, and was asked to perform medication reviews. Two pharmacists scored the agreement between ChatGPT and HCPs. The structuring of information and additional interventions suggested by ChatGPT were reviewed within an expert team. Descriptive statistics were used. Outcomes: The primary outcome was the percentage agreement between the interventions suggested by ChatGPT compared to HCPs. Secondary outcomes included the proportion of valid and incorrect interventions suggested by ChatGPT and its ability to structure patient information. Results: HCPs suggested 183 interventions and ChatGPT 202 interventions. ChatGPT achieved a 27.7% agreement with interventions of HCPs. It identified 19 additional valid interventions which HCPs missed (7.6%), but also proposed 84 incorrect interventions (33.7%). While ChatGPT demonstrated strong capability in structuring patient data (86.4% correct diagnoses linked to medication), it struggled with contextualizing appropriate laboratory target values based on patient characteristics (46.7%). Conclusion: ChatGPT had low agreement with HCPs, but found additional interventions that HCPs missed. ChatGPT lacks clinical decision-making capabilities based on individual patient contexts in older people. ChatGPT may, however, serve as a support tool to structure diagnoses and medication lists.

Original languageEnglish
Pages (from-to)1378-1385
Number of pages8
JournalJournal of the American Geriatrics Society
Volume74
Issue number5
Early online date12 Apr 2026
DOIs
Publication statusPublished - May 2026

Bibliographical note

Publisher Copyright:
© 2026 The Author(s). Journal of the American Geriatrics Society published by Wiley Periodicals LLC on behalf of The American Geriatrics Society.

Keywords

  • artificial intelligence
  • drug-related problems
  • medication therapy management

Fingerprint

Dive into the research topics of 'Conducting Medication Reviews: A Comparative Study Between ChatGPT-4 and Healthcare Professionals'. Together they form a unique fingerprint.

Cite this