Peking Union Medical College Hospital
Beijing, 100730, China
Location status: Recruiting
NCT Number: NCT07500428
This single-center, retrospective, observational study aims to construct a standardized benchmark evaluation system for intelligent breast ultrasound image interpretation and to systematically assess the diagnostic performance of current mainstream multimodal artificial intelligence (AI) models.
De-identified B-mode breast ultrasound images with confirmed pathological diagnoses will be retrospectively collected from the institutional archive (2018-2025) and supplemented with images from published open-access datasets. Expert radiologists with varying experience levels will independently annotate all images according to the American College of Radiology (ACR) Breast Imaging Reporting and Data System (BI-RADS) v2025 criteria, including glandular tissue composition, lesion characterization (mass vs. non-mass lesion), morphological descriptors, and final BI-RADS classification.
Baseline deep learning models (CNN-based ResNet-50 and Transformer-based USFM) will be trained to establish performance baselines and to stratify cases by diagnostic difficulty through cross-architecture consensus. Multiple multimodal large language models (MLLMs), including both general-purpose and medical-domain models, will then be evaluated via standardized API calls using BI-RADS-guided chain-of-thought prompts at temperature 0 for reproducibility.
Primary endpoints include BI-RADS classification accuracy and diagnostic AUC for benign-malignant differentiation. Model robustness and safety will be assessed through out-of-distribution rejection testing, temperature-stability experiments, and thinking-mode ablation studies. This study adheres to the FLAIR and TRIPOD-LLM reporting guidelines.
Interested in participating?
Request Info18 year–75 year
Female
Observational
Beijing, 100730, China
Location status: Recruiting
Background: Breast cancer is the most prevalent malignancy among women worldwide. Ultrasound is a first-line screening modality, particularly in Asian populations with dense breast tissue where mammographic sensitivity is limited. However, ultrasound interpretation is highly operator-dependent, with substantial inter-observer variability in BI-RADS classification, especially for category 4A-4B lesions. Multimodal large language models (MLLMs) have emerged as a promising tool for medical image analysis due to their zero-shot diagnostic capability, interpretable chain-of-thought reasoning, and structured report generation. Nevertheless, there is currently no standardized benchmark for evaluating AI performance in breast ultrasound interpretation.
Study Design: Approximately 1,380 breast ultrasound images will be curated (1,200 evaluation set + 150 out-of-distribution safety test set + 30 prompt development set), encompassing three diagnostic categories: normal breast, benign lesions (BI-RADS 2-4B), and malignant lesions (BI-RADS 3-5). Two junior radiologists (<5 years of experience) and two senior radiologists (>15 years) will independently annotate images per ACR BI-RADS v2025 with arbitration by a fifth expert for discordant cases.
Diagnostic difficulty will be stratified into three tiers using cross-architecture deep learning consensus: Tier 1 (straightforward, both models correct), Tier 2 (equivocal, one correct/one incorrect), and Tier 3 (difficult, both incorrect, with senior expert validation). MLLMs will be evaluated across multiple dimensions: classification accuracy, sensitivity, specificity, F1 score, AUC, Cohen's kappa agreement with expert consensus, expected calibration error (ECE), morphological feature description accuracy, and chain-of-thought reasoning quality.
Safety Assessment: (1) Out-of-distribution rejection test using 150 non-diagnostic images (degraded images, non-breast ultrasound, other imaging modalities); (2) Temperature-stability pre-experiment across parameter settings; (3) Thinking-mode ablation comparing standard vs. chain-of-thought reasoning modes. All experiments use fixed model snapshots, system fingerprint monitoring, and complete logging for reproducibility.
Healthy volunteers accepted: Yes
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
Retrospective evaluation of de-identified breast ultrasound images by multiple AI systems, including baseline deep learning models (ResNet-50, USFM) and multimodal large language models, using standardized BI-RADS-guided chain-of-thought prompts via API. No patient contact or clinical decision-making is involved.
Time frame: At study completion, approximately 12 months
Sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), and F1 score of AI models for benign-malignant classification, with histopathological diagnosis as the gold standard.
Time frame: At study completion, approximately 12 months
Overall accuracy of AI models in assigning BI-RADS categories (2, 3, 4A, 4B, 4C, 5) to breast ultrasound images, compared with expert consensus annotation as the reference standard.
Time frame: At study completion, approximately 12 months
Cohen's kappa coefficient measuring agreement between each AI model's BI-RADS classification and the expert consensus annotation, reported with 95% confidence intervals.
Time frame: At study completion, approximately 12 months
Proportion of non-diagnostic images (degraded quality, non-breast ultrasound, other imaging modalities) correctly identified and refused by AI models, evaluating domain safety.
Time frame: At study completion, approximately 12 months
Standard diagnostic performance metrics for benign-malignant classification, reported for each AI model individually.
Contact information is provided by the study sponsor or research team.
Qingli Zhu, MD
CONTACT
Yinglan Wu, MD
CONTACT
Peking Union Medical College Hospital
Other
Construction of a Standardized Benchmark Evaluation System for Intelligent Breast Ultrasound Image Interpretation and Systematic Performance Assessment of Multimodal Artificial Intelligence Models Based on ACR BI-RADS v2025 Criteria
Acronym: BUST-AI Bench
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT05386108
Brain Diseases, Brain Neoplasms
Fullerton, California, United States
View Trial DetailsNCT07581834
Breast Diseases, Breast Neoplasms
Fuzhou, Fujian, China
View Trial DetailsNCT06780176
Breast Cancer, Breast Diseases
Los Angeles, California, United States
View Trial DetailsNCT06335069
Breast Cancer, Breast Diseases
Maastricht, Netherlands
View Trial Details