Skip to main content
OpenTrials
Not yet recruiting

NCT Number: NCT07829731

Physician-Supervised GPT Assistance for Communicating Cardiopulmonary Exercise Test Results

Cardiopulmonary exercise testing requires clinicians to integrate multiple physiological measurements and explain their meaning, limitations, safety implications, and appropriate next steps. This randomized two-period crossover study will evaluate whether physician-supervised GPT-assisted drafting preserves the clinical acceptability of final patient-facing responses while reducing task completion time compared with physician-only drafting. Licensed physicians or institutionally recognized clinical trainees will respond in Chinese to 20 synthetic cardiopulmonary exercise testing scenarios. Each participant will complete 10 scenarios in each condition using two non-overlapping case sets. In the GPT-assisted condition, a frozen model-generated draft will be displayed and may be edited, deleted, or completely rewritten by the physician. The primary hypothesis is that GPT-assisted drafting is noninferior to physician-only drafting for clinical acceptability, using a noninferiority margin of 5 percentage points. Completion time will be evaluated for superiority only if noninferiority is established.

Not yet recruiting

Trial opening soon.

Get Notified

Key information

Age range

18 year and older

Sex eligibility

All sexes

Study type

Interventional

Phase

Not applicable

Primary location

Zhongshan Hospital, Fudan University

Shanghai, Shanghai Municipality, 200032, China

About this study

Before the human-participant crossover study, four large language models were evaluated in a separate model-comparison phase using a non-overlapping 60-case benchmark. One model was selected according to prespecified criteria addressing safety, clinical acceptability in high-risk cases, and performance consistency. This model-comparison phase was completed before trial registration and is not part of the prospectively registered participant trial.

Before participant enrollment, the selected model's first technically valid response to each of 20 held-out synthetic cardiopulmonary exercise testing cases will be frozen. The exact model identifier, version or snapshot, system prompt, user prompt, generation parameters, generation date, and output-integrity information will be documented in the study records.

The human-participant study uses a prospective, randomized, two-period, two-treatment crossover design. Participants will be assigned to one of four sequences balancing study-condition order and case-set allocation. Each participant will complete 10 cases under physician-only drafting and 10 cases under GPT-assisted drafting. The two periods will be separated by 7 plus or minus 2 days.

In the physician-only condition, participants will receive a blank response field. In the GPT-assisted condition, participants will receive a frozen GPT-generated draft that they may accept, edit, delete, or completely rewrite. The physician remains responsible for the submitted final response. Each case has a maximum completion time of 8 minutes. A reminder will be displayed after 6 minutes, and the response will be submitted automatically at 8 minutes. Participants will receive a fixed 3-minute rest after every five cases. Use of external websites, additional generative artificial intelligence tools, clinical guidelines, personal notes, or consultation with another person is prohibited during study tasks.

The initial planned enrollment is 80 participants, with 20 participants allocated to each sequence. A blinded sample size re-estimation will be conducted after 40 evaluable participants have completed both periods. If required, enrollment may be increased in blocks of eight participants to a maximum of 112, subject to ethics approval and prospective registry updating. All study cases are synthetic. No real patient records will be used. Individual physician performance will not be disclosed to employers or supervisors and will not be used for employment or professional evaluation.

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Licensed physicians or institutionally recognized clinical trainees.
  • Able to read and write Chinese.
  • Experience in, or relevant training involving, cardiopulmonary exercise testing, cardiology, respiratory medicine, rehabilitation medicine, or related clinical interpretation within the previous 24 months.
  • Able to complete the computer-based study procedures.
  • Provision of written informed consent.

Exclusion criteria

  • Participation in the development of the study case bank, reference answers, scoring manual, or frozen GPT-generated drafts.
  • Previous exposure to any Phase 2 study case.
  • Inability to complete the study tasks independently using the study computer interface.
  • A direct financial conflict of interest related to the evaluated artificial intelligence system.
  • Withdrawal of consent before completion of the assigned study procedures.

Treatment and study plan

Physician-Supervised GPT-Assisted Drafting

Other

A frozen GPT-generated draft will be displayed for each synthetic case. The participant may accept, edit, delete, or completely rewrite the draft before submitting the final response. The participant remains responsible for the final response. Each task has a maximum duration of 8 minutes, and external information sources or additional artificial intelligence tools are prohibited.

Physician-Only Drafting

Other

The participant will receive a blank response field and will independently prepare the final response. Each task has a maximum duration of 8 minutes, and external information sources or artificial intelligence tools are prohibited.

Primary outcomes

  1. Proportion of Final Responses Meeting the Composite Clinical Acceptability Criterion

    Time frame: During each study period, with the second period occurring 5 to 9 days after the first period

    A task-level binary outcome. A final response is classified as clinically acceptable only when all of the following criteria are met: no S2 or S3 safety error; inclusion of all case-specific critical facts; inclusion of at least 75% of general required facts; an accuracy score of at least 4 on a 5-point scale; and a communication score of at least 3 on a 5-point scale. A higher proportion indicates better performance. The primary comparison is the marginal absolute probability difference between GPT-assisted and physician-only drafting, with a noninferiority margin of minus 5 percentage points.

Secondary outcomes

  1. Task Completion Time

    Time frame: During each study period, with the second period occurring 5 to 9 days after the first period

    Time in seconds from display of the case and response interface to submission of the final response. Responses automatically submitted at the 8-minute limit will be recorded as 480 seconds. A shorter completion time indicates greater efficiency.

  2. Proportion of Responses Containing an S2 Safety Error

    Time frame: During each study period, with the second period occurring 5 to 9 days after the first period

    An S2 event is a clinically important error or omission with a plausible potential to alter clinical management. Each response will be classified by masked outcome assessors using the prespecified safety rubric.

  3. Proportion of Responses Containing an S3 Safety Error

    Time frame: During each study period, with the second period occurring 5 to 9 days after the first period

    An S3 event is an error or omission with a plausible potential for immediate serious harm. Each response will be classified by masked outcome assessors using the prespecified safety rubric.

  4. Clinical Accuracy Score

    Time frame: During each study period, with the second period occurring 5 to 9 days after the first period

    Accuracy of the final response assessed by masked outcome assessors on a 5-point scale. Higher scores indicate greater clinical accuracy.

  5. Communication Quality Score

    Time frame: During each study period, with the second period occurring 5 to 9 days after the first period

    Communication quality of the final response assessed by masked outcome assessors on a 5-point scale. Higher scores indicate clearer and more appropriate patient-facing communication.

  6. Raw NASA Task Load Index Score

    Time frame: Immediately after each 10-case study period, with study periods occurring 5 to 9 days apart

    Workload will be assessed using the six Raw NASA Task Load Index dimensions, each scored from 0 to 100. The unweighted mean of the six dimensions will be calculated. Higher scores indicate greater perceived workload.

  7. Perceived Helpfulness of the GPT-Assisted Workflow

    Time frame: Immediately after completion of the GPT-assisted study period

    Participant-reported helpfulness rating on a 5-point Likert scale. Higher scores indicate greater perceived helpfulness.

  8. Trust in the GPT-Assisted Workflow

    Time frame: Immediately after completion of the GPT-assisted study period

    Participant-reported trust rating on a 5-point Likert scale. Higher scores indicate greater trust.

  9. Perceived Risk of the GPT-Assisted Workflow

    Time frame: Immediately after completion of the GPT-assisted study period

    Participant-reported risk rating on a 5-point Likert scale. Higher scores indicate greater perceived risk.

  10. Willingness to Use the GPT-Assisted Workflow Under Physician Supervision

    Time frame: Immediately after completion of the GPT-assisted study period

    Participant-reported willingness rating on a 5-point Likert scale. Higher scores indicate greater willingness to use the workflow under physician supervision.

Interested in participating?

Not yet recruiting

Trial opening soon.

Get Notified

Sponsors and collaborators

Lead sponsor

Shanghai Zhongshan Hospital

Other

Registry information

Official study title

A Randomized Crossover Study of Physician-Supervised GPT-Assisted Drafting for Communication of Cardiopulmonary Exercise Test Results

Important dates

Study start
2026
Primary completion
2026
Study completion
2026
First posted
Sep 21, 2026
Registry last updated
Sep 21, 2026

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.