SBÜ Sultangazi Haseki Training and Research Hospital
Istanbul, Sultangazi, 34010, Turkey (Türkiye)
NCT Number: NCT07179861
This study evaluates how well anonymized artificial-intelligence (AI) tools perform on standardized pediatric case vignettes and whether showing AI suggestions can improve clinicians' answers. About 30 board-certified/eligible pediatric specialists at a single hospital complete a one-time session. Participants are randomized to two groups. Group A (n≈15): physicians answer each vignette once. Group B (n≈15): physicians answer and rate confidence (1-10), then review anonymized suggestions from five different AI tools (tool names not shown) and may keep or change their answer; changes and confidence are recorded.
Primary focus: measure AI performance (diagnostic accuracy, medication-dosing accuracy, interpretation accuracy) overall and by difficulty tier, and record AI response time. Secondary focus: quantify how AI suggestions affect human performance (change in accuracy, direction of change, confidence shift, and time). No patients or biospecimens are involved; risks are minimal (time and possible discomfort with performance review). Findings may inform safe, evidence-based ways to use AI alongside clinicians in pediatrics.
Looking for future studies?
Notify Me28 year–40 year
All sexes
Observational
Istanbul, Sultangazi, 34010, Turkey (Türkiye)
Healthy volunteers accepted: Yes
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
What: Display of AI-generated suggestions for each vignette, aggregated from five large language model tools (names not shown to participants).
When/Who: Shown only in Group 2, after the physician's initial answer and confidence score.
Purpose: Measure AI performance (primary) and quantify the effect of AI suggestions on physicians' answers (secondary).
Applies to: Group 2.
What: Self-rated confidence for the initial answer on a 1-10 scale. When/Who: Group 2 before viewing AI suggestions. Purpose: Quantify confidence changes pre- vs post-AI and relate confidence to correctness.
Applies to: Group 2.
Time frame: Day 1
Proportion of correct laboratory/imaging interpretations or appropriate next-test selections, per AI tool and pooled; stratified by difficulty tier. Unit: percent (0-100).
Time frame: Day 1
Proportion of vignettes with a correct primary diagnosis produced by each anonymized AI tool and pooled across tools. Correctness is defined against a pre-specified reference answer key; results are also stratified by pre-defined difficulty tiers (easy/moderate/difficult/very difficult). Unit of measure: percent (0-100).
Time frame: Day 1
Proportion of dose recommendations meeting pediatric standards (weight- or BSA-based ranges, route, frequency) per reference rubric, per AI tool and pooled; stratified by difficulty tier. Unit: percent (0-100).
Time frame: Day 1: Baseline (pre-AI) and immediate Post-AI within the same session (0-15 min after baseline).
Post-AI accuracy minus pre-AI accuracy per participant on the same case set; also categorized as beneficial (incorrect→correct), harmful (correct→incorrect), or no change. Accuracy is the proportion of cases with a correct final diagnosis according to a prespecified answer key.
Time frame: Day 1: Baseline (pre-AI) and immediate Post-AI within the same session (0-15 min after baseline).
Post-AI self-rated confidence minus pre-AI confidence; association with correctness is examined. Unit: scale points (-9 to +9).
Time frame: Day 1
Proportion of vignettes for which physicians revised their initial answer after AI suggestions; reported overall and by difficulty tier. Unit: percent (0-100).
Time frame: Day 1
Time from vignette display to final AI output, reported per tool and pooled; also by difficulty tier. Unit: seconds.
Time frame: Day 1
Beneficial change rate (incorrect→correct) minus harmful change rate (correct→incorrect) for diagnostic items; sensitivity analyses for dosing/interpretation. Unit: percentage points.
Haseki Training and Research Hospital
Other
A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT07352475
Artificial Intelligence (AI) in Diagnosis, Clinical Decision-making
Lille, France
View Trial DetailsNCT07068139
Artificial Intelligence (AI) in Diagnosis, Bronchial Neoplasms
View Trial DetailsNCT06936098
Artificial Intelligence (AI) in Diagnosis, Colorectal Liver Metastasis (CRLM)
Guangzhou, Guangdong, China
View Trial DetailsNCT07716670
Artificial Intelligence (AI) in Diagnosis, Biliary Disease
View Trial Details