Harvard Medical School,
Boston, Massachusetts, 02115, United States
NCT Number: NCT07741058
This study will evaluate whether artificial intelligence (AI) can enhance clinicians' accuracy, efficiency, and confidence in distinguishing lung adenocarcinoma (LUAD) from lung squamous cell carcinoma (LUSC) and kidney renal papillary cell carcinoma (KIRP) from kidney renal clear cell carcinoma (KIRC) using digitized pathology slides. These subtype classifications are routinely performed by pathologists but can be challenging and time-consuming, particularly in difficult cases.
During the study, participating clinicians will review lung and kidney pathology slides under three different conditions:
* Unaided Review: Diagnosis without AI assistance. * AI as Double-Check: The clinician first makes an independent diagnosis, after which the AI-generated diagnosis (prediction only or prediction with explanation) is revealed for review. * AI as First-Look: The AI-generated diagnosis (prediction only or prediction with explanation) is presented before the clinician begins the review.
Clinicians will be randomly assigned to different review sequences to minimize potential order effects. This study design will enable us to assess the impact of AI assistance on diagnostic accuracy, interpretation time, and clinician confidence.
Interested in participating?
Request InfoThis study aims to evaluate the effect of artificial intelligence (AI) assistance on clinicians' diagnostic performance in distinguishing lung adenocarcinoma (LUAD) from lung squamous cell carcinoma (LUSC) and kidney renal papillary cell carcinoma (KIRP) from kidney renal clear cell carcinoma (KIRC) using digitized hematoxylin and eosin (H&E)-stained whole-slide images (WSIs). ENLIGHT (Explainable Neoplasm Learning In Grounded Histology Terms) will serve as the AI system under evaluation. This is a single-session, within-reader, between-case study in which each reader evaluates distinct sets of cases under all study conditions.
The study includes three diagnostic blocks: Block X, in which WSIs are reviewed without AI assistance; Block Y1, in which clinicians make an initial diagnosis before viewing the AI output as a double-check; and Block Y2, in which the AI output is displayed before clinicians begin their review as a first-look aid. Within each AI-assisted block, the prediction-only and prediction-with-explanation sub-blocks are presented in randomized order.
Each participating pathologist will review up to 400 de-identified WSIs (up to 200 lung cancer and up to 200 kidney cancer cases). Readers will be randomly assigned to one of four study arms that differ only in the order in which Blocks X, Y1, and Y2 are completed. For each reader, distinct WSIs will be randomly assigned to the diagnostic conditions so that no WSI is reviewed more than once by the same reader.
For each case, diagnostic accuracy, time to diagnosis, and diagnostic confidence will be recorded. No reader will review the same WSI under more than one condition, thereby eliminating within-reader recall bias. In parallel, the ENLIGHT model will independently generate diagnostic predictions for all WSIs to enable direct benchmarking of AI performance against pathologists and to evaluate the impact of different AI-assisted workflows on diagnostic performance.
Healthy volunteers accepted: No
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
for Pathology Slides (i.e., Cases):
Exclusion criteria
for Pathology Slides (i.e., Cases):
Inclusion criteria
for Readers (i.e., Participants):
Readers first complete Block X (Unaided) on their assigned subset SX. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Readers first complete Block X (Unaided) on their assigned subset SX. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Readers first complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. They then complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. Then readers complete Block X (Unaided) on their assigned subset SX. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Readers first complete Block Y2 (AI as First-Look) on two separate subsets: SY2a (AI prediction-only as First-Look) and SY2b (AI prediction-with-explanation as First-Look). Within Block Y2, the order of SY2a and SY2b is randomized. They then complete Block Y1 (AI as Double-Check) on two separate subsets: SY1a (AI prediction-only as Double-Check) and SY1b (AI prediction-with-explanation as Double-Check). Within Block Y1, the order of SY1a and SY1b is randomized. Then readers complete Block X (Unaided) on their assigned subset SX. For each reader, SX, SY1a, SY1b, SY2a, and SY2b are disjoint.
Time frame: Periprocedural (at the time of slide review)
Performance of clinicians (unaided and AI-assisted) for distinguishing LUAD- LUSC and distinguishing KIRP-KIRC, measured in accuracy, sensitivity, specificity, positive predictive value, negative predictive value, and F1.
Time frame: Periprocedural (at the time of slide review)
Average time (seconds per case) required to finalize a diagnosis.
Time frame: Periprocedural (at the time of slide review)
Agreement among clinicians across conditions, measured using inter-rater reliability metrics (e.g., kappa statistics).
Time frame: Periprocedural (at the time of slide review)
The overall change in diagnostic accuracy attributable to AI assistance.
Time frame: Periprocedural (at the time of slide review)
Self-reported diagnostic confidence recorded for each case. Scale: 5 - Absolutely Certain; 4 - Mostly Certain; 3 - Unsure; 2 - Very Doubtful; 1 - Random Guess; With 5 being the highest confidence score and 1 being the lowest.
Harvard Medical School (HMS and HSDM)
Other
Acronym: ENLIGHT
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT00923065
Cancer, Neoplasms
Bethesda, Maryland, United States
View Trial DetailsNCT06412029
Cancer, Neoplasms
Newark, Delaware, United States
View Trial DetailsNCT05407844
Advanced Cancer, Cancer
Birmingham, Alabama, United States
View Trial DetailsNCT04364503
Cancer, Communicable Diseases
Los Angeles, California, United States
View Trial Details