Skip to main content
OpenTrials
Active, Not Recruiting

NCT Number: NCT07725965

Artificial Intelligence(AI) in the Emergency Department (ED) in Cologne

This retrospective, non-interventional study evaluates the prognostic performance of open-weight Large Language Models (LLMs) in the setting of a German academic emergency department. Using a full census of all consecutive emergency department cases at University Hospital Cologne between 01 January 2023 and 31 December 2025 (approximately 100,000 cases), the study assesses whether LLMs can make reliable prognostic predictions (e.g., hospital admission, imaging, diagnosis, placement) based on the initial history, vital signs, and triage category. In addition, it quantifies how strongly automated anonymization and perturbation procedures affect the models' diagnostic accuracy. This is an Investigator-Initiated Trial (IIT) with no intervention on patients.

Active, Not Recruiting

This study is active but is not currently recruiting participants.

Notify Me

Key information

Sex eligibility

All sexes

Study type

Observational

Primary location

Department of Internal Medicine II, University Hospital Cologne

Cologne, 50937, Germany

About this study

The study analyzes a retrospective cohort of all emergency department cases at the Central Emergency Department of University Hospital Cologne (01 January 2023 - 31 December 2025). Data originate from the hospital information system (HIS) and are provided in pseudonymized form via the Medical Data Integration Center (MeDIC) of University Hospital Cologne, acting as an independent trusted third party. Extracted data include sociodemographic data (age/year of birth, sex), clinical vital signs (blood pressure, heart rate, respiratory rate, oxygen saturation, temperature, GCS), medical free text (triage records, physician history and admission findings), and process/outcome data serving as the gold standard (ICD-10 diagnoses, imaging performed, admission status, timestamp/length of stay). LLM processing takes place on premise on local compute clusters or in a contractually secured enterprise environment with zero data retention; open-weight models are used. Two arms are compared: Arm A (original data) vs. Arm B (anonymized/perturbed/synthesized data). The primary endpoint is diagnostic accuracy (AUROC, F1 score) against the documented clinical outcome. Working hypotheses: (1) modern LLMs are non-inferior to the human assessment (non-inferiority); (2) modern anonymization procedures reduce model performance by less than 5% (relative performance loss).

Who can participate

Healthy volunteers accepted: No

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • All consecutive treatment cases at the Central Emergency Department of University Hospital Cologne during the period 01 January 2023 - 31 December 2025 (full census, consecutive inclusion).

Exclusion criteria

  • Documented objection to the scientific use of the data pursuant to Art. 21 General data protection Regulation (GDPR).
  • Cases lacking the minimum data required for analysis (triage/history and documented outcome).

Treatment and study plan

Primary outcomes

  1. Diagnostic accuracy of the LLM predictions (AUROC, F1 score) compared with the clinical gold standard

    Time frame: From enrollment to the end of retrospective observation period at 1 year

    assessment at the level of the individual emergency department encounter).]

Secondary outcomes

  1. Relative performance loss of the models between original data (Arm A) and anonymized/perturbed data (Arm B); hypothesis < 5%.

    Time frame: From enrollment to the end of retrospective observation period at 1 year

    Relative loss of diagnostic accuracy (AUROC, F1 score) when the models are applied to anonymized/perturbed data (Arm B) compared with original data (Arm A), expressed as the relative percentage change. Non-inferiority is assumed if the relative performance loss is below 5%.]

  2. Sensitivity, specificity, positive predictive value(PPV)/negative predictive value (NPV) for binary endpoints and agreement of the triage assessment (Cohen's kappa / Krippendorff's alpha).

    Time frame: From enrollment to the end of retrospective observation period at 1 year

    Agreement between LLM output and the documented reference for binary endpoints (e.g., admission yes/no), reported as sensitivity, specificity, and positive/negative predictive value; agreement on the ordinal triage category is reported using Cohen's kappa or Krippendorff's alpha

Sponsors and collaborators

Lead sponsor

University of Cologne

Other

Registry information

Official study title

Retrospective Validation of Large Language Models (LLM) for the Prognostic Assessment of Clinical Parameters in Emergency Department and Evaluation of the Impact of Automated Anonymisation Methods

Acronym: KINA-CO

Important dates

Study start
2026
Primary completion
2027
Study completion
2027
First posted
Jul 24, 2026
Registry last updated
Jul 28, 2026

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.