Beijing Ctiy
Beijing, Beijing Municipality, China
Location status: Recruiting
NCT Number: NCT07651280
This study will evaluate whether three-minute six-dimensions education(3M-6D education) can improve the reliability of large language models as medical assistants for the general public. Participants will be randomly assigned to receive or not receive 3M-6D education and then use ChatGPT, Gemini, or non-AI information resources. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.
Interested in participating?
Request Info18 year and older
All sexes
Interventional
Not applicable
Beijing, Beijing Municipality, China
Location status: Recruiting
This randomized, controlled, proof-of-concept simulation trial will evaluate whether three-minute six-dimensions education (3M-6D education) can improve the reliability of large language models as medical assistants for the general public.
Eligible participants will be randomly assigned in a 1:1:1:1:1 ratio to one of five study groups: the 3M-6D education GPT group, the GPT group, the 3M-6D education Gemini group, the Gemini group, or the control group. Participants in the 3M-6D education GPT and 3M-6D education Gemini groups will receive approximately three minutes of education before using ChatGPT or Gemini.Each participant will be randomly assigned one of 10 standardized clinical scenarios and complete a simulated counseling task in unrestricted natural language within approximately 10 minutes. The study will assess relevant condition identification, disposition concordance, red-flag identification, and NASA-TLX score.
Healthy volunteers accepted: Yes
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
3M-6D education is designed based on Cognitive Load Theory to reduce the cognitive burden on patients during medical interactions with AI and to improve the clarity and completeness of symptom reporting.
Guided by cognitive load theory and the natural process physicians use to take medical histories, the investigators identified candidate information dimensions and developed a structured expression framework with six dimensions for public health queries through a Delphi expert consensus process. Participants were instructed to use the framework to describe their symptoms across these six dimensions; this process can typically be completed within three minutes, so the investigators call this approach three minutes six dimensions education (3M-6D education).
Participants use ChatGPT to complete a standardized simulated clinical scenarios in unrestricted natural language.
Participants use Gemini to complete a standardized simulated clinical scenarios in unrestricted natural language.
Time frame: 1 hour.
Relevant conditions identification is defined as the proportion of participants whose final response includes the expert-defined final diagnosis or a relevant differential diagnosis.
Time frame: 1 hour.
Disposition concordance is defined as the proportion of participants whose final care recommendation matches the expert-defined level. The five levels are self-care, routine outpatient care, urgent outpatient care, emergency department visit, and emergency medical services.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Red-flag identification is defined as the proportion of participants whose final response includes the key warning signs that experts defined for the assigned scenario.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
Time frame: 1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
Time frame: 1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
Time frame: 1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
NASA-TLX score is a self-reported task-load score measured after the simulated consultation with a physician. It includes six domains: mental demand, physical demand, temporal demand, effort, frustration, and performance. Each domain is scored from 0 to 100. The total score is the mean of the six domains. Higher scores indicate greater perceived task load.
Time frame: 1 hour.
Failure to identify red flags is defined as the proportion of participants whose final response does not include the expert-defined red-flag symptoms or warning signs for the assigned standardized simulated clinical scenario.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Underestimation of disposition is defined as the proportion of participants whose final care recommendation is lower than the expert-defined disposition level for the assigned standardized simulated clinical scenario.
Time frame: 1 hour.
Time frame: 1 hour.
Time frame: 1 hour.
Contact information is provided by the study sponsor or research team.
Capital Medical University
Other
Improving the Reliability of LLMs as Medical Assistants for the General Public: a Proof of Concept Simulation Trial
Acronym: LAMP-1
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.