Skip to main content
OpenTrials
Recruiting

NCT Number: NCT07378358

Evaluation of AI Large Models for Diagnosis and Treatment in Real-World Cases: Multicenter Retrospective Study

This multicenter retrospective study aims to evaluate the diagnostic and therapeutic performance of three large language models-ChatGPT, Gemini and Deepseek-using 800 archived inpatient medical records from urology departments across four tertiary hospitals. The study will focus on the accuracy and applicability of these models in disease recognition, preliminary diagnosis and treatment recommendation generation, in order to explore their potential value and limitations in supporting clinical decision-making in real-world settings.

Recruiting

Interested in participating?

Request Info

Key information

Age range

18 year and older

Sex eligibility

All sexes

Study type

Observational

Primary location

The First Affiliated Hospital of Fujian Medical University

Fuzhou, China

Location status: Recruiting

Who can participate

Healthy volunteers accepted: No

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • The case data is sourced from the four hospitals involved in the study, with complete and authentic diagnosis and treatment records.
  • Patients must be 18 years or older, with no gender restrictions.
  • Complete medical records, including the following core information: patient' s basic information, present illness history, past medical history, physical examination, and auxiliary examinations (including laboratory and imaging tests).
  • A clear discharge diagnosis and treatment plan (including therapeutic measures and follow-up arrangements).
  • Medical records have been archived, with objective and accurate information that has not been altered.
  • The patient or their legal representative has provided informed consent, agreeing to the use of their anonymized medical data for research analysis.

Exclusion criteria

  • Medical records with significant missing information, such as key clinical details (present illness history, diagnostic or treatment records, etc.).
  • Cases where the diagnosis or treatment plan is unclear, or where treatment has not been fully completed for an initial diagnosis.
  • Cases where the primary diagnosis is not urological.
  • Cases with major errors or inconsistencies in the records that could affect further assessment.
  • Medical records in special formats or images that are not readable (e.g., handwritten notes, non-standard documentation).
  • Patients who have not signed the informed consent form or who refuse to allow their medical data to be used for research.

Treatment and study plan

Large Language Model Assessment (ChatGPT, Gemini, DeepSeek)

Other

De-identified inpatient medical records were retrospectively collected from the urology departments of four tertiary hospitals (200 cases per site, 800 in total). Each case included standardized clinical information such as demographics, chief complaint, history of present illness, past medical history, physical examination, laboratory and imaging findings, discharge diagnosis and treatment plan.

To simulate the role of an AI system in a "first-visit physician" scenario, all diagnostic conclusions, differential diagnoses and treatment plans were removed before being input into the models. Three large language models (ChatGPT, Gemini and DeepSeek) were prompted with a standardized instruction: "Based on the above clinical information, provide your preliminary diagnosis, differential diagnoses and treatment recommendations." Each model generated outputs including (i) primary and secondary diagnoses, (ii) differential diagnosis lists with reasoning and (iii) preliminary treatment suggesti

Primary outcomes

  1. Diagnostic Accuracy: Assessed by Top-1 accuracy

    Time frame: Through study completion, an average of 3 months

    Top-1: Proportion of cases where the model's first diagnosis matches the true primary diagnosis.

  2. Diagnostic Accuracy: Assessed by Top-3 accuracy

    Time frame: Through study completion, an average of 3 months

    Top-3: Proportion of cases where the true diagnosis appears in the model's top 3.

  3. Diagnostic Completeness

    Time frame: Through study completion, an average of 3 months

    Proportion of the model's diagnoses that overlap with all diagnoses (primary and secondary) in the case.

  4. Differential Diagnosis Quality

    Time frame: Through study completion, an average of 3 months

    Evaluated by experts using a Likert 5-point scale, considering factors like common disease coverage, logical clarity, and specificity

  5. Treatment Plan Quality

    Time frame: Through study completion, an average of 3 months

    Assesses whether the model's treatment suggestions align with clinical guidelines, scored by experts on completeness, appropriateness, and safety.

  6. Analysis Time

    Time frame: Through study completion, an average of 3 months

    5.Time taken by the AI model to provide diagnoses and treatment suggestions (in seconds), reflecting real-time capability.

Study contacts

Contact information is provided by the study sponsor or research team.

Sponsors and collaborators

Lead sponsor

First Affiliated Hospital of Fujian Medical University

Other

Registry information

Important dates

Study start
2026
Primary completion
2026
Study completion
2026
First posted
Jan 30, 2026
Registry last updated
Jan 30, 2026

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.