Skip to main content
OpenTrials
Recruiting

NCT Number: NCT07402668

Does AI Make Clinicians More Appropriately Confident? A Randomized Study in Preterm Birth Prediction

The goal of this randomized questionnaire-based study is to evaluate how different presentations of artificial intelligence (AI) decision support influence clinical judgment among medical doctors working in obstetrics and gynecology when assessing the risk of spontaneous preterm birth using clinical case vignettes with cervical ultrasound images. The study specifically compares two AI presentation formats: a binary classification (preterm vs term birth) and an individualized risk estimate of preterm birth.

The main questions it aims to answer are:

* Which AI presentation format leads to better alignment between clinicians' confidence and decision accuracy (diagnostic calibration)? * Do different AI presentation formats lead to helpful or harmful changes in clinical decisions?

Participants will complete an online questionnaire in which they review clinical cases, make diagnostic and management decisions, rate their diagnostic confidence before and after seeing the AI output, and report their trust in the AI.

Recruiting

Interested in participating?

Request Info

Key information

Sex eligibility

All sexes

Study type

Interventional

Phase

Not applicable

Primary location

South Jutland Hospital, Aabenraa, Denmark

Loading trial locations.

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Medical doctors currently working in or training within the field of obstetrics and gynecology.
  • Experience performing transvaginal cervical ultrasound examinations.

Exclusion criteria

  • No prior experience performing transvaginal cervical ultrasound examinations.

Treatment and study plan

AI prediction (binary)

Behavioral

AI decision support based on cervical ultrasound providing a binary classification (preterm birth before 37 weeks or term birth) in addition to standard clinical information.

AI risk estimate (%)

Behavioral

AI decision support based on cervical ultrasound providing an estimate of preterm birth risk (%) in addition to standard clinical information.

Primary outcomes

  1. Clinician diagnostic calibration (accuracy-confidence alignment) after AI exposure.

    Time frame: Immediately after AI exposure during a single questionnaire session (approximately 20 minutes).

    Agreement between post-AI decision correctness (0/1) and post-AI confidence rating (0-10) will be quantified using the Brier score. Confidence will be rescaled to 0-1 and squared differences between confidence and correctness will be averaged across cases to produce a participant-level score. Lower scores indicate better diagnostic calibration. Results will be compared between randomized arms.

Secondary outcomes

  1. Helpful switch rate and harmful switch rate.

    Time frame: Baseline (pre-AI) and immediately after AI exposure during a single questionnaire session (approximately 20 minutes).

    Proportion of cases with helpful and harmful switches calculated for each participant and compared between study arms.

    Helpful switch = incorrect pre-AI decision changing to correct post-AI decision.

    Harmful switch = correct pre-AI decision changing to incorrect post-AI decision.

  2. Change in decision accuracy, confidence, and diagnostic calibration from pre-AI to post-AI.

    Time frame: Baseline (pre-AI) and immediately after AI exposure during a single questionnaire session (approximately 20 minutes).

    Within-participant change from pre-AI to post-AI in decision accuracy (proportion of correct decisions), confidence rating, and diagnostic calibration. Differences will be compared between randomized arms and stratified by AI correctness.

  3. Association between self-rated trust in AI and behavioral reliance on AI.

    Time frame: Immediately after AI exposure during a single questionnaire session (approximately 20 minutes).

    Self-rated trust in the AI output will be measured using a numeric rating scale (0-10) after AI exposure for each case. Behavioral reliance will be quantified as the proportion of post-AI decisions concordant with the AI output. The relationship between trust ratings and behavioral reliance, including concordance when the AI is correct and incorrect, will be evaluated at the participant level and compared between randomized arms.

  4. Follow-up cervical ultrasound planning.

    Time frame: Baseline (pre-AI) and immediately after AI exposure during a single questionnaire session (approximately 20 minutes).

    Proportion of cases in which clinicians plan an additional cervical ultrasound (yes/no), summarized per participant and compared pre-post AI and between randomized arms.

Study contacts

Contact information is provided by the study sponsor or research team.

Emilie Pi F Sejer, MD

CONTACT

[email protected]

0045 28890690

Sponsors and collaborators

Lead sponsor

Rigshospitalet, Denmark

Other

Collaborators

  • Copenhagen Academy for Medical Education and Simulation
  • Department of Computer Science, University of Copenhagen, Denmark
  • Technical University of Denmark
  • The Foundation of 17.12.1981

Registry information

Important dates

Study start
2026
Primary completion
2026
Study completion
2026
First posted
Feb 11, 2026
Registry last updated
Jul 1, 2026

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.