Our objective is to make the early diagnosis and assessment of AD and MCI based on multimodal deep learning. Initially, Gait disorder, facial expression identification dysfunction, and speech and language impairment are of great significance in the occurrence and development of AD and MCI. However, due to the high complexity and stealthiness of these clinical symptoms, no uniform conclusions have been made. Hence, we attempt to apply machine learning methods to recognize the video and audio information. In this way, we will explore the changing characteristics of gait, expression, and language in AD and MCI, and analyze their diagnostic effectiveness as diagnostic markers, finally providing new ideas and experimental data for the diagnosis of AD and MCI. Secondly, multimodal medical information needs to be integrated and comprehensively analyzed. We aim to propose an optimal diagnostic strategy referring to the different degrees of dependency on multimodal medical information in diagnosis. Moreover, observing the changes in multimodal medical information with the progress of AD and MCI, we expect to build a predicting model of AD diagnosis and prognosis.
The methods are as follows:
- Collecting multimodal medical information A variety of multimodal medical information would be carefully collected including the baseline demographic data, chief complaint and medical history, peripheral organ function assessment, laboratory examination, imaging examination, neuroelectrophysiological examination, neurocognitive and psychological examination, information on gait, expression, and language, and biological samples, etc.
- Revealing the changes of gait, expression, and language in patients with AD and MCI, and verifying their diagnostic efficacy.
For multimodal medical information on gait, OpenPose model was used to extract human key points and construct a human skeleton structure diagram. Based on graph neural networks and convolutional neural networks, instantaneous action analysis of single-frame images is carried out. And then utilizing the Transformer model, gait sequence analysis is carried out by integrating multi-frame video.For multimodal medical information on facial expression, the Dlib algorithm will be used to extract facial key points, combined with facial expression images, and the spatiotemporal Transformer model will be used for facial expression analysis. For multimodal medical information on language, ASRT model will be used for speech recognition and text content extraction. Simultaneously, the frequency domain Fourier transform and wavelet transform will be applied to extract frequency domain information and analyze the speech features by integrating language content, voice intonation, speech speed, and other information. Based on the attention model, the gait, expression, and language analysis results of AD and MCI will be compared with those of the control group to reveal the features of AD and MCI and provide evidence for disease diagnosis.
- Analyzing the different degrees of dependency on multimodal information in the diagnosis of AD and MCI diseases, and establishing an optimal diagnosis strategy In the supervised learning process, the attention mechanism-based method will be used to analyze the influence of multimodal information on the final results. At the same time, based on the knowledge map, the patient's blood biochemical indicators, genomic information and other fields of knowledge would be added to the model. Based on Bayesian probability inference and causal inference theory, the causal programming method will be used to model the causal analysis of information and diagnosis results of different modes. Based on AutoML method, multimodal information will be combined and optimized, and a reliable optimal diagnosis strategy will be established according to experimental results.
- Exploring the changes of multimodal medical information with the progression of the disease, and build a predicting model for early diagnosis and disease progression of AD.
Viewing multimodal medical information as the control condition, the Transformer model will be used to model time sequence information, and the conditional diffusion model will be used to generate patients' MRI image changes and other disease progression-related information, providing the basis for disease progression prediction. Based on the large multimodal model technology, the output of the model will be interfered with and adjusted referring to the judgment and description of professional doctors, to generate the prediction in line with the judgment of professional doctors, and finally construct the interpretable early diagnosis and disease progression prediction model.