University of Pennsylvania Health System
Philadelphia, Pennsylvania, 19101, United States
NCT Number: NCT04574882
This project seeks to identify and characterize features derived from digital data (e.g. social media, online search, mobile media) which are associated with coronary heart disease (CHD) and related risk factors, and develop models that use digital data and conventional predictive models to predict CHD risk and health care utilization.
Looking for future studies?
Notify Me30 year–74 year
All sexes
Observational
Philadelphia, Pennsylvania, 19101, United States
Cardiovascular disease is the leading cause of death in the US. While secondary prevention approaches have improved longevity of patients, risk factors and adverse health behaviors (e.g., physical inactivity, smoking) are highly prevalent, and in most contemporary series, less than 1% of adults meet all factors of ideal CV health. The logistics and practicalities of meeting the goal of ideal CV health have not been clearly elucidated. Practice guidelines recommend using the Framingham risk score (FRS) or other risk prediction tools to classify patients' risk of CV disease. These models however are imprecise and there is increasing focus on identifying markers that provide better measures of risk. As digital platforms are increasingly used to document lifestyle and health behaviors, data from digital sources may provide a window into manifestations of novel risk factors and potentially a better characterization of existing risk factors. While it seems like a cliche to mention the profound impact of digital data on everyday lives, there is indeed great substance in the opportunities these new media provide for understanding behavioral, social, and environmental determinants of health. This project seeks to identify and characterize features derived from digital data (e.g. social media, online search, mobile media) which are associated with coronary heart disease (CHD) and related risk factors, and develop models that use digital data and conventional predictive models to predict CHD risk and health care utilization.
Healthy volunteers accepted: Yes
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
Interested participants may complete the informed consent online. After informed consent, the participant will be asked to share the digital data types that they use (Facebook, Instagram, Twitter, Google search, step data) and then participants will complete a cross-sectional survey.
Time frame: Through study completion, an average of 3 years
The primary outcome is topics and features (derived using the LDA method for clustering language data).
For each participant, we included all available Facebook wall posts from the start of their account history through data collection, regardless of whether they occurred before or after a CHD diagnosis. We examined associations between linguistic features (unigrams, LIWC categories, LDA topics) and cardiovascular case status (CHD presence vs absence) using Pearson correlation and logistic regression. Latent LDA, a systematic method to identify text-based themes, was applied to generate 200 clusters of co-occurring words ("topics"). For each feature type (unigram, LIWC category, LDA topic), we fit separate logistic regression models and calculated Pearson correlation coefficients to assess predictive value for case status. Each language-derived feature was encoded as a normalized frequency count per user to enable consistent comparison across participants.
Time frame: Through study completion, an average of 3 years
Reliability in predicting CHD related event in patient as measured by Framingham Risk Score.
The Framingham Risk Score (FRS) is a validated means of predicting cardiovascular disease (CVD) risk. Input variables include age, cigarette smoking, total cholesterol, HDL cholesterol, systolic blood pressure measurement and treatment for hypertension. Point values are calculated based on each of these risks. A 10-year risk score can be derived as a percentage. Risk scores range from 0-20%.
Low Risk: Less than 10% risk that you will develop a heart attack or die from coronary disease in the next 10 years.
Intermediate risk: A 10 to 20% risk that you will develop a heart attack or die from coronary disease in the next 10 years.
High Risk: A greater than 20% risk that you will develop a heart attack or die from coronary disease in the next 10 years.
Time frame: Through study completion, an average of 3 years
Prediction of cost for health care utilization between heart disease and non- heart disease subjects measured by insurance claims data
University of Pennsylvania
Other
Using Digital Data to Predict Cardiovascular Health and Health Care Utilization
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT03188705
Arterial Occlusive Diseases, Arteriosclerosis
Lancaster, Pennsylvania, United States
View Trial DetailsNCT05882045
Body Weight, Cardiovascular Diseases
Phoenix, Arizona, United States
View Trial DetailsNCT02439775
Cardiovascular Diseases, Hypertension
Huntsville, Alabama, United States
View Trial DetailsNCT05820295
Arrhythmias, Cardiac, Atrial Fibrillation
New York, United States
View Trial Details