Department of Internal Medicine II, University Hospital Cologne
Cologne, 50937, Germany
NCT Number: NCT07725965
This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.
This study is active but is not currently recruiting participants.
Notify MeAll sexes
Observational
Cologne, 50937, Germany
The study analyzes a retrospective cohort of all emergency department cases at the Central Emergency Department of University Hospital Cologne (01 January 2023 - 31 December 2025). Data originate from the hospital information system (HIS) and are provided in pseudonymized form via the Medical Data Integration Center (MeDIC) of University Hospital Cologne, acting as an independent trusted third party. Extracted data include sociodemographic data (age/year of birth, sex), clinical vital signs (blood pressure, heart rate, respiratory rate, oxygen saturation, temperature, GCS), medical free text (triage records, physician history and admission findings), and process/outcome data serving as the gold standard (ICD-10 diagnoses, imaging performed, admission status, timestamp/length of stay). LLM processing takes place on premise on local compute clusters or in a contractually secured enterprise environment with zero data retention; open-weight models are used. Two arms are compared: Arm A (original data) vs. Arm B (anonymized/perturbed/synthesized data). The primary endpoint is diagnostic accuracy (AUROC, F1 score) against the documented clinical outcome. Working hypotheses: (1) modern LLMs are non-inferior to the human assessment (non-inferiority); (2) modern anonymization procedures reduce model performance by less than 5% (relative performance loss).
Healthy volunteers accepted: No
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
Time frame: From enrollment to the end of retrospective observation period at 1 year
assessment at the level of the individual emergency department encounter).]
Time frame: From enrollment to the end of retrospective observation period at 1 year
Relative loss of diagnostic accuracy (AUROC, F1 score) when the models are applied to anonymized/perturbed data (Arm B) compared with original data (Arm A), expressed as the relative percentage change. Non-inferiority is assumed if the relative performance loss is below 5%.]
Time frame: From enrollment to the end of retrospective observation period at 1 year
Agreement between LLM output and the documented reference for binary endpoints (e.g., admission yes/no), reported as sensitivity, specificity, and positive/negative predictive value; agreement on the ordinal triage category is reported using Cohen's kappa or Krippendorff's alpha
University of Cologne
Other
Retrospective Validation of Large Language Models (LLM) for the Prognostic Assessment of Clinical Parameters in Emergency Department and Evaluation of the Impact of Automated Anonymisation Methods
Acronym: KINA-CO
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT07676968
Emergency Care, POCUS
View Trial DetailsNCT07546071
Emergency Care, Emergency Department Crowding
Ramat Gan, Israel
View Trial DetailsNCT07248605
Emergency Care, Obstetric
Al Mansurah, Egypt
View Trial DetailsNCT07257705
Clinical Decision Support Systems, Diarrhea Infectious
Kalongo, Agago District, Uganda
View Trial Details