Skip to main content
OpenTrials
Completed

NCT Number: NCT07179861

Comparing Artificial Intelligence and Physicians: A Vignette-Based Study in Pediatric Clinical Decision-Making

This study evaluates how well anonymized artificial-intelligence (AI) tools perform on standardized pediatric case vignettes and whether showing AI suggestions can improve clinicians' answers. About 30 board-certified/eligible pediatric specialists at a single hospital complete a one-time session. Participants are randomized to two groups. Group A (n≈15): physicians answer each vignette once. Group B (n≈15): physicians answer and rate confidence (1-10), then review anonymized suggestions from five different AI tools (tool names not shown) and may keep or change their answer; changes and confidence are recorded.

Primary focus: measure AI performance (diagnostic accuracy, medication-dosing accuracy, interpretation accuracy) overall and by difficulty tier, and record AI response time. Secondary focus: quantify how AI suggestions affect human performance (change in accuracy, direction of change, confidence shift, and time). No patients or biospecimens are involved; risks are minimal (time and possible discomfort with performance review). Findings may inform safe, evidence-based ways to use AI alongside clinicians in pediatrics.

Completed

Looking for future studies?

Notify Me

Key information

Age range

28 year–40 year

Sex eligibility

All sexes

Study type

Observational

Primary location

SBÜ Sultangazi Haseki Training and Research Hospital

Istanbul, Sultangazi, 34010, Turkey (Türkiye)

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Board-certified or board-eligible pediatric specialist (general pediatrics) (in the first 10 years of expertise)
  • Actively practicing at the participating institution/network at the time of enrollment.
  • Able and willing to complete all vignette items individually in a single session and to follow study instructions for the assigned cohort (direct answers or confidence rating + viewing anonymized AI suggestions).
  • Fluent in Turkish and able to use a computer interface.
  • Provides written informed consent.

Exclusion criteria

  • Pediatric subspecialist practice as primary role (e.g., cardiology, infectious diseases, neurology, neonatology, etc.), to maintain a homogeneous general pediatrics cohort.
  • Prior access to or participation in creating the study vignettes, answer keys, or scoring rubrics; direct involvement with the study team.
  • Inability to complete the session without external help or use of non-protocol resources (internet/AI tools) during answering (outside of anonymized AI suggestions shown by the system in Group 2).
  • Failure to complete ≥90% of items or major protocol deviation (e.g., discussion with others during the task).
  • Any condition judged by investigators to interfere with valid participation (e.g., severe time constraints, inability to provide consent).

Treatment and study plan

AI Suggestions (Anonymized 5-tool panel)

Other

What: Display of AI-generated suggestions for each vignette, aggregated from five large language model tools (names not shown to participants).

When/Who: Shown only in Group 2, after the physician's initial answer and confidence score.

Purpose: Measure AI performance (primary) and quantify the effect of AI suggestions on physicians' answers (secondary).

Applies to: Group 2.

Confidence Rating Task (1-10 Likert)

Other

What: Self-rated confidence for the initial answer on a 1-10 scale. When/Who: Group 2 before viewing AI suggestions. Purpose: Quantify confidence changes pre- vs post-AI and relate confidence to correctness.

Applies to: Group 2.

Primary outcomes

  1. AI Interpretation Accuracy (%)

    Time frame: Day 1

    Proportion of correct laboratory/imaging interpretations or appropriate next-test selections, per AI tool and pooled; stratified by difficulty tier. Unit: percent (0-100).

  2. AI Diagnostic Accuracy (%)

    Time frame: Day 1

    Proportion of vignettes with a correct primary diagnosis produced by each anonymized AI tool and pooled across tools. Correctness is defined against a pre-specified reference answer key; results are also stratified by pre-defined difficulty tiers (easy/moderate/difficult/very difficult). Unit of measure: percent (0-100).

  3. AI Medication-Dosing Accuracy (%)

    Time frame: Day 1

    Proportion of dose recommendations meeting pediatric standards (weight- or BSA-based ranges, route, frequency) per reference rubric, per AI tool and pooled; stratified by difficulty tier. Unit: percent (0-100).

Secondary outcomes

  1. Change in Physician Diagnostic Accuracy (percentage points) (Group 2 only)

    Time frame: Day 1: Baseline (pre-AI) and immediate Post-AI within the same session (0-15 min after baseline).

    Post-AI accuracy minus pre-AI accuracy per participant on the same case set; also categorized as beneficial (incorrect→correct), harmful (correct→incorrect), or no change. Accuracy is the proportion of cases with a correct final diagnosis according to a prespecified answer key.

  2. Confidence Shift (Δ on a 1-10 scale) (Group 2 only)

    Time frame: Day 1: Baseline (pre-AI) and immediate Post-AI within the same session (0-15 min after baseline).

    Post-AI self-rated confidence minus pre-AI confidence; association with correctness is examined. Unit: scale points (-9 to +9).

  3. Answer-Change Frequency (%) (Group 2 only)

    Time frame: Day 1

    Proportion of vignettes for which physicians revised their initial answer after AI suggestions; reported overall and by difficulty tier. Unit: percent (0-100).

  4. AI Response Time (seconds per vignette)

    Time frame: Day 1

    Time from vignette display to final AI output, reported per tool and pooled; also by difficulty tier. Unit: seconds.

  5. Net Benefit Index of AI Exposure (percentage points) (Group 2 only)

    Time frame: Day 1

    Beneficial change rate (incorrect→correct) minus harmful change rate (correct→incorrect) for diagnostic items; sensitivity analyses for dosing/interpretation. Unit: percentage points.

Sponsors and collaborators

Lead sponsor

Haseki Training and Research Hospital

Other

Registry information

Official study title

A Prospective, Cross-Sectional, Vignette-Based Observational Study Comparing Clinical Decision-Making Performance of Pediatriciansand AI Models

Important dates

Study start
2025
Primary completion
2025
Study completion
2025
First posted
Sep 18, 2025
Registry last updated
Sep 23, 2025

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.