Skip to main content
OpenTrials
Not Yet Recruiting

NCT Number: NCT07045207

Evaluating AI and Human Expert Decisions in Colorectal Cancer

The goal of this observational study is to evaluate the decision-making consistency between large language models (LLMs) and expert multidisciplinary teams (MDTs) in adult patients diagnosed with colorectal cancer who underwent MDT consultation between January 2023 and December 2024.

The main questions it aims to answer are:

How consistent are the treatment decisions generated by LLMs compared to actual MDT decisions? Do different LLMs (e.g., ChatGPT, DeepSeek) show varying levels of agreement with expert recommendations? What clinical factors contribute to differences between AI-generated and human expert decisions? Researchers will compare the AI-generated treatment recommendations with real-world MDT decisions using anonymized patient records to see if LLMs can reliably support clinical decision-making in oncology.

Participants will:

Have their de-identified clinical data (e.g., imaging, pathology, MDT notes) processed through several LLMs Not be contacted or receive any interventions, as this is a retrospective study using existing clinical records only.

Not Yet Recruiting

Trial opening soon.

Get Notified

Key information

Sex eligibility

All sexes

Study type

Observational

Primary location

About this study

This is a retrospective, non-interventional observational study aiming to evaluate the consistency between treatment decisions made by large language models (LLMs) and multidisciplinary team (MDT) experts in the management of colorectal cancer (CRC).

Colorectal cancer is a highly heterogeneous malignancy requiring personalized treatment strategies, often developed through MDT discussions that integrate input from surgery, oncology, radiology, pathology, and other specialties. While MDTs improve treatment planning and outcomes, they are time- and resource-intensive, and subject to variability in expert judgment. With the rise of artificial intelligence, especially LLMs such as ChatGPT and DeepSeek, there is growing interest in their potential role in assisting or standardizing clinical decision-making.

In this study, researchers will retrospectively analyze de-identified clinical records of approximately 1,500 patients with histologically confirmed colorectal cancer who underwent MDT consultation at a tertiary cancer center between January 2023 and December 2024. Key clinical data-including demographic information, imaging reports (CT, MRI), endoscopy results, pathology findings, and MDT recommendations-will be extracted and anonymized.

These de-identified records will be input into several LLMs (ChatGPT, DeepSeek, Baichuan, and Qwen) running on secure offline servers. The models will be asked to generate treatment recommendations, which will be categorized into predefined decision codes (e.g., surgery, systemic therapy, chemoradiotherapy, further diagnostics). Each case will be input three times to assess the consistency of the model output.

The primary outcome is the agreement between AI-generated recommendations and original MDT decisions, quantified using Cohen's Kappa. Secondary analyses include comparison among LLMs using chi-squared tests, evaluation of output consistency via Fleiss' Kappa, and identification of clinical factors associated with discordant decisions.

This study does not involve any direct patient contact, intervention, or new clinical procedures. All data are historical and anonymized in accordance with ethical and legal requirements. The results are expected to inform the potential value, limitations, and appropriate use of AI in supporting multidisciplinary decision-making in oncology.

Who can participate

Healthy volunteers accepted: No

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Patients with a histologically confirmed diagnosis of colorectal cancer
  • Patients who received multidisciplinary team (MDT) consultation at Peking University Cancer Hospital between January 1, 2023 and December 31, 2024
  • Availability of complete clinical records, including(MDT consultation notes, CT or MRI imaging reports, Pathology reports, Outpatient or inpatient medical summaries)

Exclusion criteria

  • Incomplete or missing medical records related to MDT decision-making
  • MDT consultations conducted for non-oncologic purposes (e.g., hernia evaluation, stoma planning)
  • Missing critical clinical data such as imaging or pathology reports
  • Duplicate or conflicting records that prevent reliable data analysis

Treatment and study plan

LLM-MDT

Other

Leveraging large language models (LLMs) to Generate Multidisciplinary Team (MDT) Treatment Recommendations

Primary outcomes

  1. Agreement Between AI-Generated and MDT Treatment Decisions

    Time frame: January 1, 2023 to December 31, 2024 (based on MDT consultation date)

    Description: The primary outcome is the consistency between treatment recommendations generated by large language models (LLMs) and those made by expert multidisciplinary teams (MDTs) for colorectal cancer cases. Consistency will be quantified using Cohen's Kappa coefficient. Higher Kappa values indicate stronger agreement

Secondary outcomes

  1. Comparison of Agreement Across Different AI Models

    Time frame: January 1, 2023 to December 31, 2024

    Description: To compare the consistency of treatment decisions generated by different large language models (e.g., ChatGPT, DeepSeek, Baichuan, Qwen) with expert MDT decisions using Cohen's Kappa and chi-squared tests. This outcome assesses whether performance varies across AI models.

  2. Output Stability of AI Models on Repeated InputsDescription

    Time frame: January 1, 2023 to December 31, 2024

    To evaluate the reproducibility of treatment decisions generated by each AI model when the same clinical case is input multiple times. Stability will be assessed using Fleiss' Kappa to measure consistency across repeated outputs.

  3. Identification of Clinical Factors Associated With Decision Discordance

    Time frame: January 1, 2023 to December 31, 2024

    To identify key clinical features (e.g., disease stage, metastasis status, treatment history) that are associated with discordant treatment decisions between AI models and MDT experts. Statistical analysis will be conducted to explore which case characteristics lead to lower agreement.

Study contacts

Contact information is provided by the study sponsor or research team.

Yongjiu Chen, PhD

CONTACT

[email protected]

+86 18813041827

Sponsors and collaborators

Lead sponsor

Peking University Cancer Hospital & Institute

Other

Registry information

Official study title

Comparison of Large Language Models and Expert Multidisciplinary Team Decisions in Colorectal Cancer

Important dates

Study start
2025
Primary completion
2026
Study completion
2026
First posted
Jul 1, 2025
Registry last updated
Jul 1, 2025

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.