Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study

Background
Large language model (LLM)–based artificial intelligence chatbots direct the power of large training data sets toward successive, related tasks as opposed to single-ask tasks, for which artificial intelligence already achieves impressive performance. The capacity of LLMs to assist in the full scope of iterative clinical reasoning via successive prompting, in effect acting as artificial physicians, has not yet been evaluated.

Objective
This study aimed to evaluate ChatGPT’s capacity for ongoing clinical decision support via its performance on standardized clinical vignettes.

Methods
We inputted all 36 published clinical vignettes from the Merck Sharpe & Dohme (MSD) Clinical Manual into ChatGPT and compared its accuracy on differential diagnoses, diagnostic testing, final diagnosis, and management based on patient age, gender, and case acuity. Accuracy was measured by the proportion of correct responses to the questions posed within the clinical vignettes tested, as calculated by human scorers. We further conducted linear regression to assess the contributing factors toward ChatGPT’s performance on clinical tasks.

Results
ChatGPT achieved an overall accuracy of 71.7% (95% CI 69.3%-74.1%) across all 36 clinical vignettes. The LLM demonstrated the highest performance in making a final diagnosis with an accuracy of 76.9% (95% CI 67.8%-86.1%) and the lowest performance in generating an initial differential diagnosis with an accuracy of 60.3% (95% CI 54.2%-66.6%). Compared to answering questions about general medical knowledge, ChatGPT demonstrated inferior performance on differential diagnosis (β=–15.8%; P<.001) and clinical management (β=–7.4%; P=.02) question types.

Conclusions
ChatGPT achieves impressive accuracy in clinical decision-making, with increasing strength as it gains more clinical information at its disposal. In particular, ChatGPT demonstrates the greatest accuracy in tasks of final diagnosis as compared to initial diagnosis. Limitations include possible model hallucinations and the unclear composition of ChatGPT’s training data set.

Topics

No keywords indexed for this article. Browse by subject →

References

29

[1]

Artificial intelligence in healthcare

Kun-Hsing Yu, Andrew L. Beam, Isaac S. Kohane

Nature Biomedical Engineering 10.1038/s41551-018-0305-z

[2]

10.2196/27850

[3]

10.1016/j.jacr.2021.01.013

[4]

10.1038/s41598-022-24721-5

[5]

10.1097/md.0000000000029587

[6]

10.1038/s41467-022-29437-8

[7]

10.1259/bjro.20210062

[8]

10.30953/bhty.v4.176

[9]

ChatGPT: optimizing language models for dialogueOpen AI202211302023-02-15https://openai.com/blog/chatgpt/

[10]

10.1371/journal.pdig.0000198

[11]

10.2139/ssrn.4314839

[12]

10.2139/ssrn.4335905

[13]

10.2139/ssrn.4322372

[14]

TerwieschCWould Chat GPT3 get a Wharton MBA? a prediction based on its performance in the operations management courseMack Institute for Innovation Management at the Wharton School, University of Pennsylvania20232023-08-02https://mackinstitute.wharton.upenn.edu/wp-content/uploads/2023/01/Christian-Terwiesch-Chat-GTP.pdf

[15]

10.1001/jama.2023.1344

[16]

10.1038/s41746-021-00423-6

[17]

10.1016/j.jacr.2023.05.003

[18]

10.1101/2023.01.30.23285067

[19]

10.1101/2023.02.02.23285399

[20]

10.1101/2023.02.21.23285886

[21]

Case studiesMerck Manual, Professional Version2023-02-01https://www.merckmanuals.com/professional/pages-with-widgets/case-studies?mode=list

[22]

10.1016/j.jemermed.2008.03.038

[23]

10.1016/j.jopan.2021.03.009

[24]

Artificial intelligence and algorithmic bias: implications for health systems

Trishan Panch, Heather Mattie, Rifat Atun

Journal of Global Health 10.7189/jogh.09.020318

[25]

Smedley, BD Unequal Treatment: Confronting Racial and Ethnic Disparities in Health Care (2003)

[26]

10.1145/3368555.3384448

[27]

10.1145/3442188.3445922

[28]

Survey of Hallucination in Natural Language Generation

Ziwei Ji, Nayeon Lee, Rita Frieske et al.

ACM Computing Surveys 10.1145/3571730

[29]

10.48550/arxiv.2211.09527

Cited By

282

Letter to editor on "Integrating expert knowledge into large language models improves performance for psychiatric reasoning and diagnosis"

Hooman Hadianfard, Reza Moshfeghinia · 2026

Psychiatry Research

AI-assisted learning: ChatGPT for anamnesis in women’s health education

Daniela S Espírito Santo, Thiago Lott Bezerra · 2026

Medical Education Online

Exploring the accuracy of embedded ChatGPT-4 and ChatGPT-4o in generating BI-RADS scores: a pilot study in radiologic clinical support

Dan Nguyen, Arya Rao · 2025

Clinical Imaging

A Future of Self-Directed Patient Internet Research: Large Language Model-Based Tools Versus Standard Search Engines

Arya Rao, Andrew Mu · 2025

Annals of Biomedical Engineering

Development and Evaluation of an Artificial Intelligence–Powered Surgical Oral Examination Simulator: A Pilot Study

Arya Rao, Siona Prasad · 2025

Mayo Clinic Proceedings: Digital He...

The Use of an Artificial Intelligence Platform OpenEvidence to Augment Clinical Decision-Making for Primary Care Physicians

Ryan T. Hurt, Christopher R. Stephenson · 2025

Journal of Primary Care & Commu...

Large Language Models for Chatbot Health Advice Studies

Bright Huo, Amy Boyle · 2025

JAMA Network Open

Efficacy and empathy of AI chatbots in answering frequently asked questions on oral oncology

Rata Rokhshad, Zaid H. Khoury · 2025

Oral Surgery, Oral Medicine, Oral P...

Evaluating large language models and agents in healthcare: key challenges in clinical applications

Xiaolan Chen, Jiayang Xiang · 2025

Intelligent Medicine

Dedicated AI Expert System vs Generative AI With Large Language Model for Clinical Diagnoses

Mitchell J. Feldman, Edward P. Hoffer · 2025

JAMA Network Open

Large Language Models in Neurological Practice: Real-World Study

Natale Vincenzo Maiorana, Sara Marceglia · 2025

Journal of Medical Internet Researc...

Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial

Emilia Brügge, Sarah Ricchizzi · 2024

BMC Medical Education

Large Language Models Can Enable Inductive Thematic Analysis of a Social Media Corpus in a Single Prompt: Human Validation Study

Michael S Deiner, Vlad Honcharov · 2024

JMIR Infodemiology

Security Implications of AI Chatbots in Health Care

Jingquan Li · 2023

Journal of Medical Internet Researc...

Metrics

282

Citations

29

References

Details

Published: Aug 22, 2023
Vol/Issue: 25
Pages: e48659

Authors

Cite This Article

Arya Rao, Michael Pang, John Kim, et al. (2023). Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study. Journal of Medical Internet Research, 25, e48659. https://doi.org/10.2196/48659

Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow: Development and Usability Study

You May Also Like