Beijing Tsinghua Changgung Hospital
Beijing, Changping, 102218, China
Location status: Recruiting
NCT Number: NCT07538882
The precise treatment of primary hepatocellular carcinoma (HCC) highly depends on accurate disease staging (CNLC, TNM, BCLC) and scientific treatment decision-making, which necessitate the integration of both imaging and clinical baseline data. This study prospectively recruits HCC patients and clinical physicians across different hospital tiers to evaluate the clinical value of a self-developed artificial intelligence (AI) model in assisting multi-dimensional comprehensive assessment and treatment decision-making. Utilizing a Multi-Rater Multi-Case (MRMC) crossover balanced design, the study compares the accuracy of clinical evaluations performed by physicians under "unassisted (without AI)" versus "AI-assisted" conditions. A key focus is to explore whether AI can significantly enhance the comprehensive assessment capabilities of physicians in primary/secondary care hospitals, thereby prospectively reducing diagnostic and therapeutic heterogeneity across different institutional levels.
Interested in participating?
Request Info18 year and older
All sexes
Observational
Beijing, Changping, 102218, China
Location status: Recruiting
Brief Summary: The precise treatment of primary hepatocellular carcinoma (HCC) highly depends on accurate disease staging (CNLC, TNM, BCLC) and scientific treatment decision-making, which necessitate the integration of both imaging and clinical baseline data. This study prospectively recruits HCC patients and clinical physicians across different hospital tiers to evaluate the clinical value of a self-developed artificial intelligence (AI) model in assisting multi-dimensional comprehensive assessment and treatment decision-making. Utilizing a Multi-Rater Multi-Case (MRMC) crossover balanced design, the study compares the accuracy of clinical evaluations performed by physicians under "unassisted (without AI)" versus "AI-assisted" conditions. A key focus is to explore whether AI can significantly enhance the comprehensive assessment capabilities of physicians in primary/secondary care hospitals, thereby prospectively reducing diagnostic and therapeutic heterogeneity across different institutional levels.
Gold Standard (Reference Standard): The reference standard (Ground Truth) for all prospectively enrolled cases is established by an independent expert panel consisting of 3 authoritative experts. The panel determines the final standard answers for the four classification tasks through blinded independent evaluation and joint discussion (voting system), incorporating complete prospective imaging data, clinical baseline data, multidisciplinary team (MDT) consensus, and final pathological or clinical follow-up results.
2.1 Evaluator Eligibility:
2.2 Patient/Case Eligibility:
Inclusion criteria
Exclusion criteria
Intervention Model: Crossover Assignment Masking: Single Blind. Participating evaluators are blinded to the gold standard answers of the cases and to the evaluation results of other participating physicians.
Arms and Interventions:
Case Set Partition: 108 prospectively and consecutively enrolled eligible HCC cases are batched and randomly divided into Dataset Set A (54 cases) and Dataset Set B (54 cases). It is ensured that there are no statistically significant differences between the two sets regarding tumor burden, liver function grading, and staging distribution.
Evaluator Grouping: A total of 12 prospectively recruited clinical physicians are included, comprising 4 in the tertiary hospital senior group, 4 in the tertiary hospital junior group, and 4 in the primary/secondary hospital group. They are divided into two evaluation groups based on stratified randomization:
Group A (6 evaluators): 2 tertiary senior, 2 tertiary junior, 2 primary/secondary hospital.
Group B (6 evaluators): 2 tertiary senior, 2 tertiary junior, 2 primary/secondary hospital.
Arm 1 - Group A Evaluators:
Phase 1 Intervention (Control): Independent evaluation of Set A (54 cases) combining clinical texts and imaging data, recording 4 classification results, without AI assistance.
Phase 2 Intervention (Experimental): Evaluation of Set B (54 cases). The system presents the AI model's 4 prediction results and related evidence; physicians provide the final judgment after comprehensive reference.
Arm 2 - Group B Evaluators:
Phase 1 Intervention (Control): Independent evaluation of Set B (54 cases) combining clinical texts and imaging data, recording 4 classification results, without AI assistance.
Phase 2 Intervention (Experimental): Evaluation of Set A (54 cases). Physicians provide the final judgment after referencing the AI model's results.
Primary Outcome:
Improvement in Overall Accuracy: The difference in average accuracy across the 4 classification tasks between AI-assisted evaluation (experimental group) and independent evaluation (control group).
Secondary Outcomes:
Homogenization Effect: Assessment of whether the difference in clinical evaluation accuracy between physicians in the primary/secondary hospital group and the tertiary hospital groups is significantly reduced under AI assistance.
Evaluation Efficiency: Comparison of the average evaluation time per case between physicians with and without AI assistance.
Inter-rater Agreement: Comparison of the consistency of evaluation results among physicians (e.g., using Kappa statistics), with and without AI assistance.
Sample Size Justification:
The sample size calculation for this study is based on the expected change in the overall average accuracy across all levels of prospectively recruited physicians. It is estimated that the overall average accuracy without AI assistance is 0.60, and with AI assistance is 0.70.
Setting the significance level for a two-sided test at 0.05 (corresponding to a Z-value of approximately 1.96) and the statistical power at 0.80 (corresponding to a Z-value of approximately 0.84), the sample size was determined using the standard statistical method for comparing two independent proportions. Assuming no clustering effect resulting from multiple case evaluations by the same physician, this calculation indicates that each intervention group requires at least 353 independent evaluations.
Power Verification:
In the actual configuration of this study, there are 12 physicians in total.
Total independent evaluations for the control group (without AI) = Group A (6 evaluators) x Set A (54 cases) + Group B (6 evaluators) x Set B (54 cases) = 648 independent evaluations.
Total independent evaluations for the experimental group (with AI) also = 648 independent evaluations.
Since 648 evaluations is greater than the required base of 353 evaluations, the current configuration of cases and physicians already possesses sufficient statistical power. This sample size provides a conservative margin (approximately 1.8 times the base requirement) to adequately account for any clustering effect (intra-class correlation) resulting from multiple case evaluations by the same physician in this MRMC design.
Healthy volunteers accepted: No
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
Physicians independently evaluate the HCC cases and provide staging and treatment decisions using only complete clinical baseline data and imaging data, without any assistance from the AI model.
Physicians evaluate the HCC cases and provide final staging and treatment decisions after reviewing the initial predictions and related evidence generated by the self-developed artificial intelligence (AI) model, alongside the clinical baseline and imaging data.
Time frame: Up to 1 week (Assessed upon completion of all case evaluations)
The difference in average accuracy across the 4 classification tasks between AI-assisted evaluation (experimental) and independent evaluation (control). Accuracy is determined by comparing physicians' predictions against the reference standard (Ground Truth) established by the independent expert panel
Time frame: Up to 1 week (Assessed upon completion of all case evaluations)
Assessment of whether the difference in clinical evaluation accuracy between physicians in the primary/secondary hospital group and the tertiary hospital groups is significantly reduced under AI assistance compared to unassisted independent evaluation.
Time frame: Up to 1 week (Assessed upon completion of all case evaluations)
Comparison of the average evaluation time (e.g., measured in minutes) per case required by participating physicians when utilizing AI assistance versus performing unassisted independent evaluation.
Time frame: Up to 1 week (Assessed upon completion of all case evaluations)
Comparison of the consistency of evaluation results (staging and treatment decisions) among all participating physicians, assessed using appropriate statistical measures (e.g., Kappa statistics), under AI-assisted versus unassisted conditions.
Contact information is provided by the study sponsor or research team.
Beijing Tsinghua Chang Gung Hospital
Other
A Prospective, Randomized, Controlled, Crossover Study of Artificial Intelligence-Assisted Multi-Dimensional Staging and Treatment Decision-Making for Hepatocellular Carcinoma
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT07291076
Adenocarcinoma, Carcinoma
Duarte, California, United States
View Trial DetailsNCT07419841
Adenocarcinoma, Carcinoma
Boston, Massachusetts, United States
View Trial DetailsNCT07224750
Adenocarcinoma, Carcinoma
Duarte, California, United States
View Trial DetailsNCT04634357
Adenocarcinoma, Carcinoma
Los Angeles, California, United States
View Trial Details