Skip to main content
OpenTrials
Active, Not Recruiting

NCT Number: NCT07744555

Change in the Proportion of Correct Interpretations of Pulmonary Opacities

This randomized crossover clinical trial will evaluate whether an artificial intelligence-based expert system improves the interpretation of chest radiographs by final-year medical students, rural physicians, and general practitioners with less than two years of clinical experience.

Participants will interpret chest radiographs containing normal findings or pulmonary opacities classified as alveolar, interstitial, or mixed. Each participant will review the same set of radiographs twice: once without assistance from the expert system and once with assistance from the artificial intelligence system. The order of these two reading conditions will be randomly assigned, with a six-week interval between sessions to reduce memory effects.

The main outcome will be the proportion of correct interpretations compared with a previously established reference standard based on radiologist interpretation supported by chest computed tomography findings. The study will also assess diagnostic confidence and agreement with the reference standard.

This study does not involve treatment decisions or direct patient care. It is intended to determine whether artificial intelligence support may improve the accuracy and confidence of less-experienced physicians when interpreting pulmonary opacities on chest radiographs.

Active, Not Recruiting

This study is active but is not currently recruiting participants.

Notify Me

Key information

Age range

18 year and older

Sex eligibility

All sexes

Study type

Interventional

Phase

Not applicable

Primary location

Clínica Universidad de La Sabana

Chía, Cundinamarca, 111321, Colombia

About this study

Study Design

A randomized, controlled, two-period crossover trial with a six-week washout period will be conducted to evaluate the change in the proportion of correct interpretations of pulmonary opacities on chest radiographs by final-year medical students, physicians in mandatory rural or social service, and general practitioners with less than two years of clinical experience, before and after assistance from an artificial intelligence-based image-recognition expert system.

Each participant serves as their own control, allowing direct comparison of interpretation performance with and without assistance from the expert system. Each participant evaluates a set of chest radiographs under two conditions: Phase A, interpretation without expert-system assistance, and Phase B, interpretation with expert-system assistance. The order of the phases is randomized. Sequence 1 begins with Phase A and, after a six-week washout period, proceeds to Phase B; Sequence 2 begins with Phase B and, after the washout period, proceeds to Phase A. The six-week interval is considered sufficient to reduce memory or learning effects. During this period participants have no access to the previously evaluated images and receive no feedback about their responses.

Expert System

The expert system is a freely available, open-source deep-learning model for computer-aided interpretation of chest radiographs that reports the presence of pulmonary opacities. Its output is displayed together with each image during Phase B only. The system is used exclusively as a reading aid for study participants; it is not used for clinical care, and no diagnostic, treatment or management decision for any patient is based on its output.

Study Population

Participants are final-year undergraduate medical students, physicians completing mandatory rural or social service, and licensed general practitioners with less than two years of clinical experience who agree to participate as chest radiograph readers. Eligible participants are randomly selected from the available list and invited individually to participate. Participants with corrected visual acuity worse than 20/60 (0.3 on the decimal scale) in the better-seeing eye are excluded, because such a limitation in resolving fine detail may be a source of performance bias when interpreting digital radiographs on calibrated monitors, independently of the intervention. Visual acuity is verified by documentation of acuity with corrective lenses, when applicable, or by a simple visual acuity test before the reading session.

Participant-level variables collected are sex, prior training in radiology or radiograph interpretation, and previous rotation or work experience in radiology services. Image-level variables are the type of pulmonary opacity according to Fleischner Society terminology (alveolar, interstitial or mixed), laterality, and severity of pulmonary involvement. The study is conducted at Universidad de La Sabana and Clinica Universidad de La Sabana.

Reference Standard

The reference standard was established in a previous study conducted by the same research group and is not generated in the present trial. In that study, two board-certified radiologists independently interpreted each chest radiograph under blinded conditions, without knowledge of the other reader's findings or of the expert-system output. Radiographs classified as normal and radiographs with abnormal findings were included. When the two readers disagreed, the case was reviewed by a third board-certified radiologist who examined the radiograph together with the corresponding chest computed tomography examination; that reading defines the final reference standard for the case.

In the present protocol the assigned radiologist does not redefine the reference standard. The radiologist only verifies that the images used correspond to the previously classified database and supervises the technical allocation of cases, without modifying the original diagnoses. All comparisons involving the expert system and the participants are therefore made against a single, previously established diagnosis.

All radiographs are reviewed by the supervising radiologist to confirm uniform technical quality, including adequate penetration, inspiration and centering; technically suboptimal images are excluded. Only cases in which the interval between chest radiography and chest computed tomography is no longer than seven days are selected, in order to reduce bias from disease progression and to ensure stability of the opacity pattern.

Primary Outcome Definition

The primary outcome is dichotomous: correct or incorrect interpretation, defined hierarchically. An interpretation is classified as globally correct only when both criteria are met in comparison with the reference standard: the participant correctly identifies whether a pulmonary opacity is present or absent, and, when an opacity is present, correctly classifies the pattern as alveolar, interstitial or mixed. Secondary analyses evaluate separately the diagnostic accuracy of opacity detection (sensitivity and specificity), the accuracy of pattern classification, diagnostic confidence, and agreement with the reference standard.

Sample Size

The sample size was calculated from the expected proportion of pulmonary opacity detection with and without the expert system, assuming a two-sided significance level of 0.05, 90% power, a minimum clinically relevant difference of 5 percentage points and a within-subject standard deviation of 10 percentage points. The required sample size was 47 participants; allowing for 10% loss during data collection, the final sample size is 50 participants, each serving as their own control.

Each participant completes 16 interpretations without expert-system assistance and 16 interpretations with assistance, for a total of 32 readings. The image database contains 410 chest radiographs from hospitalized patients. For each participant the image set is balanced and includes 4 radiographs with an alveolar pattern, 4 with an interstitial pattern, 4 with a mixed pattern and 4 normal radiographs, so that every participant evaluates the different patterns under comparable conditions.

Data Collection

Chest radiographs and chest computed tomography examinations from hospitalized patients are evaluated retrospectively. Eligible cases are patients who underwent both studies within a maximum interval of 15 days, between January 2019 and May 2024. Data are collected from records available in the hospital database of Universidad de La Sabana using an electronic form designed for this study, which includes participants' sociodemographic variables and the variables required for classification of pulmonary opacities. Data collection is supervised by a general practitioner with experience in radiology, who trains participants on completion of the form.

Immediately after recording each interpretation, participants rate their confidence regarding the presence or absence of a pulmonary opacity on a 5-point ordinal Likert scale: 1, none (pure guess); 2, low (less than 50% confidence); 3, moderate (50-70%); 4, high (70-90%); 5, very high (90-100%).

Before the sessions begin, the correct functioning of the electronic form and of the expert-system software is verified. Participants receive an explanation of the study protocol and of the measures adopted to protect data confidentiality, and are blinded to the clinical diagnoses of the cases. One reading session is conducted per participant in each phase. Images are interpreted in a quiet, distraction-free room with ambient lighting appropriate for radiological interpretation, on calibrated monitors with adequate resolution and brightness. Each participant uses the same workstation, and mobile devices are not permitted during the sessions.

Bias Control

Observer bias is minimized by interpreting radiographs under blinded conditions, without access to clinical information or chest computed tomography findings. Selection bias is reduced by random sampling of eligible participants and by the crossover design, in which each participant serves as their own control. Verification bias is controlled by comparing all interpretations with the previously established reference standard; no new reference standard is generated and the images are not reinterpreted during the study. Interobserver variability is reduced through standardized training of participants and evaluators, and randomization of the image presentation order reduces bias related to learning effects or visual fatigue.

Randomization and Allocation Concealment

An independent investigator, not involved in recruitment or assessment, generates the allocation sequence for the 50 participants using R software version 2025.09.1+401; set.seed() ensures reproducibility and the sample() function assigns participants to Sequence 1 or Sequence 2.

Allocation is concealed using sequentially numbered, opaque, sealed envelopes. The same independent investigator prepares 50 identical envelopes, each containing a card with the assigned sequence. Envelopes are opaque, sealed and consecutively numbered from 1 to 50, and are kept by the principal investigator responsible for data management. Allocation occurs only after the participant has met all eligibility criteria and signed the informed consent form; the participant's name is written on the next consecutively numbered envelope, which is then opened to irreversibly reveal the assigned sequence. This procedure prevents the recruiter and the participant from knowing the assigned sequence in advance.

Each participant evaluates exactly the same chest radiographs in both phases, so that diagnostic difficulty is identical under both conditions and any observed difference can be attributed to the expert system. The order of image presentation within each phase is randomized independently for each participant using the sample() function in R.

Statistical Analysis

Data are analyzed with Stata version 17 (StataCorp LLC, College Station, TX, USA). A descriptive analysis of all variables is performed first. Categorical variables are summarized with absolute and relative frequencies; continuous variables are presented as mean and standard deviation when normally distributed and as median and interquartile range otherwise, with normality assessed by the Shapiro-Wilk test.

Because each participant serves as their own control, paired tests are used. McNemar's test compares, between the phase without and the phase with expert-system assistance, the proportion of globally correct interpretations, the proportion of correct detections and the proportion of correct pattern classifications. The Wilcoxon signed-rank test compares median diagnostic confidence scores between phases.

Percentage agreement between participant interpretations and the reference standard is calculated, and Cohen's kappa is estimated to assess diagnostic reproducibility, interpreted according to Landis and Koch as poor (<0.20), fair (0.21-0.40), moderate (0.41-0.60), substantial (0.61-0.80) and almost perfect (>0.80). Effect size is expressed as differences in proportions with 95% confidence intervals. Prevalence-adjusted and bias-adjusted kappa estimates are also calculated.

To explore carryover effects, a conditional logistic regression model including the order of exposure as a covariate is used; this model also allows adjustment for prior experience in radiograph interpretation, sex and level of medical training. Statistical significance is set at a two-sided p-value below 0.05. All analyses follow the intention-to-treat principle, including all participants assigned to the study regardless of whether they complete both phases.

Results will be reported in accordance with the CONSORT extension for randomized crossover trials, the STARD guidelines for diagnostic accuracy studies, and the GRRAS guidelines for reliability and agreement studies.

Who can participate

Healthy volunteers accepted: Yes

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Final-year undergraduate medical students.
  • Physicians completing mandatory rural or social service.
  • Licensed general practitioners with less than two years of clinical experience.
  • Willingness to participate as a chest radiograph reader.
  • Ability to complete both study reading sessions.

Exclusion criteria

  • Corrected visual acuity worse than 20/60 in the better-seeing eye.
  • Visual-field restriction, impaired contrast perception, binocular vision abnormalities, or another visual condition that prevents adequate interpretation of chest radiographs.
  • Inability to complete the visual acuity assessment before the reading session.

Treatment and study plan

Artificial intelligence expert system

Device

Artificial intelligence-based expert system for computer-aided interpretation of chest radiographs. The system is an open-source deep-learning model that analyzes posteroanterior chest radiographs and reports the presence of pulmonary opacities. During the intervention phase, each participant interprets the same set of chest radiographs while the output of the expert system is displayed together with each image. Participants remain responsible for the final interpretation and record it on the study data form. The expert system is used only as a reading aid for study participants; it is not used for clinical care, and no diagnostic, treatment, or management decision for any patient is based on its output.

Primary outcomes

  1. Proportion of globally correct interpretations of pulmonary opacities

    Time frame: Each of the two reading sessions, separated by a 6-week washout period

    Proportion of chest radiographs interpreted correctly by the participant, compared with a previously established reference standard. An interpretation is counted as globally correct only when both criteria are met: (1) the participant correctly identifies whether a pulmonary opacity is present or absent, and (2) when an opacity is present, the participant correctly classifies the pattern as alveolar, interstitial or mixed. The outcome is expressed as the percentage of correct interpretations out of the 16 radiographs read in each condition, and is compared between the reading condition without expert-system assistance and the reading condition with expert-system assistance.

Secondary outcomes

  1. Sensitivity and specificity for the detection of pulmonary opacities

    Time frame: Each of the two reading sessions, separated by a 6-week washout period

    Diagnostic accuracy of participants for detecting the presence or absence of a pulmonary opacity on chest radiographs, expressed as sensitivity and specificity with 95% confidence intervals, using the previously established reference standard as the comparator. Values are calculated separately for the reading condition without expert-system assistance and the reading condition with expert-system assistance.

  2. Proportion of correctly classified pulmonary opacity patterns

    Time frame: Each of the two reading sessions, separated by a 6-week washout period

    Among radiographs with a pulmonary opacity according to the reference standard, the proportion in which the participant correctly classifies the pattern as alveolar, interstitial or mixed, according to Fleischner Society terminology. Values are calculated separately for the reading condition without expert-system assistance and the reading condition with expert-system assistance.

  3. Diagnostic confidence score

    Time frame: Each of the two reading sessions, separated by a 6-week washout period

    Self-reported confidence of the participant in the interpretation of each chest radiograph regarding the presence or absence of a pulmonary opacity, rated immediately after each reading on a 5-point ordinal Likert scale where 1 = none (pure guess), 2 = low, 3 = moderate, 4 = high and 5 = very high. Median scores are compared between the reading condition without expert-system assistance and the reading condition with expert-system assistance.

  4. Agreement between participant interpretation and the reference standard (Cohen's kappa)

    Time frame: Each of the two reading sessions, separated by a 6-week washout period

    Level of agreement between the participant's interpretation and the previously established reference standard, estimated with Cohen's kappa coefficient. Kappa values are interpreted according to the Landis and Koch criteria as poor (<0.20), fair (0.21-0.40), moderate (0.41-0.60), substantial (0.61-0.80) and almost perfect (>0.80). Prevalence-adjusted and bias-adjusted kappa estimates are also calculated. Values are reported separately for the reading condition without expert-system assistance and the reading condition with expert-system assistance.

Sponsors and collaborators

Lead sponsor

Universidad de la Sabana

Other

Registry information

Official study title

Change in the Proportion of Correct Interpretations of Pulmonary Opacities by General Practitioners Using an Expert System

Acronym: INTERPRET

Important dates

Study start
2026
Primary completion
2026
Study completion
2026
First posted
Aug 4, 2026
Registry last updated
Aug 4, 2026

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.