A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study

Friederike Holderried; Christian Stegemann-Philipps; Anne Herrmann-Werner; Teresa Festl-Wietek; Martin Holderried; Carsten Eickhoff; Moritz Mahling

doi:10.2196/59213

journal article Aug 16, 2024

A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study

Friederike Holderried

Christian Stegemann-Philipps

JMIR Medical Education Vol. 10 pp. e59213 · JMIR Publications Inc.

View at Publisher Save 10.2196/59213

Abstract

Background
Although history taking is fundamental for diagnosing medical conditions, teaching and providing feedback on the skill can be challenging due to resource constraints. Virtual simulated patients and web-based chatbots have thus emerged as educational tools, with recent advancements in artificial intelligence (AI) such as large language models (LLMs) enhancing their realism and potential to provide feedback.

Objective
In our study, we aimed to evaluate the effectiveness of a Generative Pretrained Transformer (GPT) 4 model to provide structured feedback on medical students’ performance in history taking with a simulated patient.

Methods
We conducted a prospective study involving medical students performing history taking with a GPT-powered chatbot. To that end, we designed a chatbot to simulate patients’ responses and provide immediate feedback on the comprehensiveness of the students’ history taking. Students’ interactions with the chatbot were analyzed, and feedback from the chatbot was compared with feedback from a human rater. We measured interrater reliability and performed a descriptive analysis to assess the quality of feedback.

Results
Most of the study’s participants were in their third year of medical school. A total of 1894 question-answer pairs from 106 conversations were included in our analysis. GPT-4’s role-play and responses were medically plausible in more than 99% of cases. Interrater reliability between GPT-4 and the human rater showed “almost perfect” agreement (Cohen κ=0.832). Less agreement (κ<0.6) detected for 8 out of 45 feedback categories highlighted topics about which the model’s assessments were overly specific or diverged from human judgement.

Conclusions
The GPT model was effective in providing structured feedback on history-taking dialogs provided by medical students. Although we unraveled some limitations regarding the specificity of feedback for certain feedback categories, the overall high agreement with human raters suggests that LLMs can be a valuable tool for medical education. Our findings, thus, advocate the careful integration of AI-driven feedback mechanisms in medical training and highlight important aspects when LLMs are used in that context.

Topics

No keywords indexed for this article. Browse by subject →

References

47

[1]

10.1136/bmj.2.5969.486

[2]

Peterson, MC West J Med (1992)

[3]

10.1046/j.1525-1497.1999.00267.x

[4]

10.1186/1472-6920-12-16

[5]

10.1016/j.pec.2005.06.004

[6]

10.1016/j.pec.2018.04.013

[7]

10.1186/s12909-023-04533-5

[8]

10.3109/0142159x.2011.531170

[9]

10.1111/medu.13387

[10]

A scoping review: virtual patients for communication skills in medical undergraduates

Síle Kelly, Erica Smyth, Paul Murphy et al.

BMC Medical Education 10.1186/s12909-022-03474-9

[11]

The effectiveness of using virtual patient educational tools to improve medical students’ clinical reasoning skills: a systematic review

Ruth Plackett, Angelos P. Kassianos, Sophie Mylan et al.

BMC Medical Education 10.1186/s12909-022-03410-x

[12]

Artificial Intelligence Supporting the Training of Communication Skills in the Education of Health Care Professions: Scoping Review

Tjorven Stamer, Jost Steinhäuser, Kristina Flägel

Journal of Medical Internet Research 10.2196/43311

[13]

10.2196/53961

[14]

10.1111/medu.14152

[15]

10.1007/s10586-018-2334-5

[16]

Proof of Concept: Using ChatGPT to Teach Emergency Physicians How to Break Bad News

Jeremy J Webb

Cureus 10.7759/cureus.38755

[17]

10.1080/01421590600622665

[18]

10.1097/acm.0000000000001578

[19]

Practical and ethical challenges of large language models in education: A systematic scoping review

Lixiang Yan, Lele Sha, Linxuan Zhao et al.

British Journal of Educational Technology 10.1111/bjet.13370

[20]

10.1016/j.tsc.2023.101440

[21]

Utilizing OpenAI's GPT‐4 for written feedback

Makenna Carlson, Austin Pack, Juan Escalante

TESOL Journal 10.1002/tesj.759

[22]

10.1056/aioa2400196

[23]

10.1016/j.caeai.2021.100027

[24]

10.1148/radiol.230582

[25]

How Does ChatGPT Perform on the United States Medical Licensing Examination (USMLE)? The Implications of Large Language Models for Medical Education and Knowledge Assessment

Aidan Gilson, Conrad W Safranek, Thomas Huang et al.

JMIR Medical Education 10.2196/45312

[26]

10.1227/neu.0000000000002632

[27]

10.2196/52113

[28]

Survey of Hallucination in Natural Language Generation

Ziwei Ji, Nayeon Lee, Rita Frieske et al.

ACM Computing Surveys 10.1145/3571730

[29]

OpenAI Platform2024-02-03https://platform.openai.com

[30]

10.48550/arxiv.2402.05359

[31]

R: A Language and Environment for Statistical Computing (2023)

[32]

The Measurement of Observer Agreement for Categorical Data

J. Richard Landis, Gary G. Koch

Biometrics 1977 10.2307/2529310

[33]

Bias, prevalence and kappa

Ted Byrt, Janet Bishop, John B. Carlin

Journal of Clinical Epidemiology 10.1016/0895-4356(93)90018-v

[34]

10.1016/j.chbah.2023.100022

[35]

FagbohunOHarrisonRMDereventsovAAn empirical categorization of prompting techniques for large language models: a practitioner's guidearXiv20242024-03-25http://arxiv.org/abs/2402.14837

[36]

GPT-42024-03-25https://openai.com/research/gpt-4

[37]

10.48550/arxiv.2201.11903

[38]

10.1109/acii59096.2023.10388213

[39]

The Role of ChatGPT, Generative Language Models, and Artificial Intelligence in Medical Education: A Conversation With ChatGPT and a Call for Papers

Gunther Eysenbach

JMIR Medical Education 10.2196/46885

[40]

10.1109/icalt58122.2023.00100

[41]

10.1186/s41239-023-00425-2

[42]

10.1016/j.caeai.2023.100199

[43]

10.3390/nu16070914

[44]

HäkkinenJRamadanZA Study on the Perception of Feedback with Varying Sentiment Generated Using a Large Language Model20232024-03-25Stockholm, Swedemhttps://www.diva-portal.org/smash/get/diva2:1779789/FULLTEXT01.pdf

[45]

10.1007/s10212-022-00654-5

[46]

Bowman, SR arXiv (2023)

[47]

10.1016/j.compedu.2023.104967

Cited By

94

Large Language Model-Based Virtual Patient Simulations in Medical and Nursing Education: A Review

Young-Woo Jo, Myungeun Lee · 2025

Applied Sciences

Large Language Models for Information Retrieval: Challenges and Chances

Timo Breuer, Sameh Frihat · 2025

Datenbank-Spektrum

Large language models improve clinical decision making of medical students through patient simulation and structured feedback: a randomized controlled trial

Emilia Brügge, Sarah Ricchizzi · 2024

BMC Medical Education

Metrics

94

Citations

47

References

Details

Published: Aug 16, 2024
Vol/Issue: 10
Pages: e59213

Authors

F

Friederike Holderried

C

Christian Stegemann-Philipps

Cite This Article

Friederike Holderried, Christian Stegemann-Philipps, Anne Herrmann-Werner, et al. (2024). A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study. JMIR Medical Education, 10, e59213. https://doi.org/10.2196/59213

A Language Model–Powered Simulated Patient With Automated Feedback for History Taking: Prospective Study

You May Also Like