Skip to main content
OpenTrials
Completed

NCT Number: NCT06208423

Physician Reasoning on Management Cases With Large Language Models

This study will evaluate the effect of providing access to GPT-4, a large language model, compared to traditional management decision support tools on performance on case-based management reasoning tasks.

Completed

Looking for future studies?

Notify Me

Key information

Sex eligibility

All sexes

Study type

Interventional

Phase

Not applicable

Primary location

Stanford University

Palo Alto, California, 94304, United States

About this study

Artificial intelligence (AI) technologies, specifically advanced large language models like OpenAI's ChatGPT, have the potential to improve medical decision-making. Although ChatGPT-4 was not developed for its use in medical-specific applications, it has demonstrated promise in various healthcare contexts, including medical note-writing, addressing patient inquiries, and facilitating medical consultation. However, little is known about how ChatGPT augments the clinical reasoning abilities of clinicians.

Clinical reasoning is a complex process involving pattern recognition, knowledge application, and probabilistic reasoning. Integrating AI tools like ChatGPT-4 into physician workflows could potentially help reduce clinician workload and decrease the likelihood of mismanagement. However, ChatGPT-4 was not developed for clinical reasoning nor has it been validated for this purpose. Further, it may be subject to disinformation, including convincing confabulations that may mislead clinicians. If clinicians misuse this tool, it may not improve reasoning and could even cause harm. Therefore, it is important to study how clinicians use large language models to augment clinical reasoning prior to routine incorporation into patient care.

In this study, participants will be randomized to answer clinical management cases with or without access to ChatGPT-4. Each case has multiple components, and the participants will be asked to discuss their reasoning for each component. Answers will be graded by independent reviewers blinded to treatment assignment. A grading rubric was developed for each case by a panel of 4-7 expert discussants. Discussants independently developed a rubric for each case, and then any discrepancies were resolved through multiple rounds of discussions.

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Participants must be licensed physicians and have completed at least post-graduate year 2 (PGY2) of medical training.
  • Training in Internal medicine, family medicine, or emergency medicine.

Exclusion criteria

  • Not currently practicing clinically.

Treatment and study plan

GPT-4

Other

OpenAI's GPT-4 large language model with chat interface.

Primary outcomes

  1. Management Reasoning

    Time frame: Within one-hour study

    Percent correct (range: 0 to 100) for each case.

Secondary outcomes

  1. Time Spent on Management

    Time frame: Within one-hour study

    Time (in minutes) participants spend per case between the two study arms.

Sponsors and collaborators

Lead sponsor

Stanford University

Other

Collaborators

  • Beth Israel Deaconess Medical Center
  • University of Minnesota

Registry information

Official study title

Management Reasoning With AI Chat Bots

Important dates

Study start
2023
Primary completion
2024
Study completion
2024
First posted
Jan 17, 2024
Registry last updated
Sep 27, 2024

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.