Skip to main content
OpenTrials
Completed

NCT Number: NCT04432961

Natural Language Processing (NLP) Analysis of Free Text Notes to Investigate Coronavirus (COVID-19)

A retrospective cohort study investigating clinical notes using Natural Language Processing in combination with structured data from the Electronic Health Record (EHR) to create a database for analytics to identify features associated with outcomes.

Completed

Looking for future studies?

Notify Me

Key information

Age range

18 year–100 year

Sex eligibility

All sexes

Study type

Observational

Primary location

Cambridge University NHS Foundation Trust

Cambridge, United Kingdom

About this study

Patients admitted to Cambridge University Hospitals (CUH)with COVID-19 have undergone routine clinical documentation and specific investigation and testing for COVID-19. The pathway for these patients ranges from supportive measures on the ward to deterioration requiring Intensive therapy Unit (ITU) admission and ventilatory support. Patients are also at risk of developing complications such as Acute Kidney Injury and thromboembolism. Identification of the risk factors for these and other outcomes such as the requirement for ventilation remain a challenge and reviewing the clinical data for these patients is critical in the understanding of the relationship between patient characteristics and outcomes.

There is data available in structured fields in the EHR, however, this is sometimes incomplete and inaccurate. An assessment of the free text clinical notes provides an opportunity to fill in the gaps and provide a much richer dataset for evaluation. We plan to use Natural Language Processing (NLP) (a field of machine learning that allows computers to analyse human language) to review Discharge Summaries of patients admitted to hospital with COVID-19 and convert free text data into structured data for analysis.

The NLP techniques developed by Dr Collier's team include methods for coding of free texts to SNOMED CT and other biomedical ontologies. These methods, based on statistical machine learning from human annotated texts, have been benchmarked for scientific texts and social media. In this project we intend to adapt these techniques for patient records. The techniques will require a number of human annotated patient records in order to adapt. The NLP output will be combined with structured data from the EHR and undergo statistical analysis to identify the rates of complications in patients with COVID-19 and risk factors associated with these. This may help to guide management decisions by earlier intervention to prevent poor outcomes in these patients.

Who can participate

Only the study team can determine whether someone qualifies for participation.

Inclusion criteria

  • Male and female
  • Age range: 18 to 100 years
  • Patients admitted to Cambridge University Hospitals with confirmed COVID-19 on lab testing

Exclusion criteria

Children and patients with a negative COVID test.

Treatment and study plan

Primary outcomes

  1. research database of EHR records from COVID-19 patients processed using NLP tools for named entity recognition and linking adapted to CUH EMR data to identify variables of interest

    Time frame: 1 year

    Our overarching hypothesis is that the NLP-extracted data from the free-text discharge summary can be combined with structured data from the EMR to yield insights into the development of complications. Patient with severe disease requiring ITU admission and non severe disease managed on an inpatient ward will be included. The variables of interest will include patient characteristics and specific encounter related information including length of stay and baseline investigations (e.g., blood tests) and interventions received

Secondary outcomes

  1. A set of annotation guidelines to produce human-expert (gold) labelled data for a subset of the EHR

    Time frame: 6 months

  2. A comparison of the NLP output to terms in the structured problem list to identify missing terms in the structured problem list

    Time frame: 1 year

Sponsors and collaborators

Lead sponsor

Cambridge University Hospitals NHS Foundation Trust

Other

Collaborators

  • University of Cambridge

Registry information

Official study title

A Database and Analytics Study of Free Text Clinical Notes and Structured Data to Investigate Phenotype Associations With Outcomes in Patients With COVID-19

Important dates

Study start
2020
Primary completion
2021
Study completion
2021
First posted
Jun 16, 2020
Registry last updated
Jul 28, 2021

OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.

View the official ClinicalTrials.gov record (opens in a new tab)

This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.

Published trials that share one or more normalized conditions with this study.