Beijing Friendship Hospital, Capital Medical University
Beijing, China
Location status: Recruiting
NCT Number: NCT07795645
The goal of this observational study is to leverage the abundant patient resources and standardized medical records from Beijing Friendship Hospital, Xuanwu Hospital, and Beijing Anzhen Hospital, combined with the existing data and knowledge platform of guidelines, consensus, medical literature, and dialogue data from Beijing Haitian Ruisheng Science Technology Co.,Ltd, with Beijing Zhilan Medical Technology Co., Ltd. conducting the fine-tuning, optimization, and validation of the medical large language model. The model is fine-tuned according to the consultation and diagnostic needs of different departments to improve the quality and efficiency of hospital medical services, enhance intelligence, and elevate the level of medical care. Through deployment to hospitals at all levels, it aims to achieve standardized services and support graded diagnosis and treatment. The overall research includes medical big data construction, medical knowledge graph construction, medical large model training and fine-tuning, and large model application platform development and deployment.
Interested in participating?
Request Info18 year and older
All sexes
Observational
Beijing, China
Location status: Recruiting
This is a multicenter, ambispective cohort study conducted across three hospitals in Beijing (Beijing Friendship Hospital as the lead site, Xuanwu Hospital, and Beijing Anzhen Hospital), utilizing approximately 80000 retrospective historical medical records from patients diagnosed with chronic gastritis, gastric cancer, gastro esophageal reflux, coronary artery disease, or stroke treated between July 2014 and June 2024, along with approximately 16000 prospectively enrolled patients with the same diseases between July 2024 and June 2027.
Healthy volunteers accepted: No
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
(1) Age ≥ 18 years; (2) Diagnosed with any of the following: chronic gastritis, gastric cancer, gastro esophageal reflux, coronary artery disease, or stroke.
Exclusion criteria
(1) Records with information that cannot be correctly read due to modification or smudging; (2) Examination reports that are smudged or damaged, making them uninterpretable by the large model.
(1) Severe psychiatric disorders (e.g., depression, mania, epilepsy, schizophrenia); (2) Judged by the investigator to be unable to comply with study procedures; (3) Poor audio quality due to accent or recording issues that prevents accurate data capture; (4) Laboratory or imaging reports that are smudged or damaged, making them uninterpretable by the model
Prospective Historical Medical Records:
Integrating multimodal medical data (electronic medical records, medical images and reports, laboratory results, genetic information) with medical guidelines, expert consensus, clinical databases, medical literature, encyclopedias, patents, and doctor-patient dialogue data. Datasets are constructed in phases: pre-training general knowledge datasets (guidelines, textbooks, literature, historical records) and fine-tuning disease-specific datasets (doctor-patient dialogues, disease-specific records, treatment plans, follow-up records). Techniques include data cleaning, standardization, transformation, annotation, and augmentation. The platform adopts a human-machine collaboration strategy to reduce data annotation costs, combining medical experts' professional knowledge with AI capabilities.
Extracting multimodal knowledge graphs from guidelines, consensus, and literature using the QLora training framework. The process involves medical knowledge modeling (defining entities, relations, attributes), entity recognition (disease names, drug names), relation extraction (disease-symptom, drug-treatment relationships), attribute extraction (incidence, dosage), and knowledge fusion and completion to build a tens-of-millions-level knowledge graph. The construction incorporates multimodal data integration (text, images, audio) through image recognition and speech recognition technologies.
Developing fine-tuning techniques for different disease areas using the ChatGLM large model, with integration of expert feedback. Pre-training uses large-scale medical datasets with language modeling tasks. Fine-tuning incorporates physician annotation, data augmentation (multi-task learning, adversarial training), and prompt-tuning to generate disease-specific responses. Deployment applies quantization to reduce model size and improve inference speed. A quality evaluation system is developed as a standardized method for measuring model performance, incorporating customized methods tailored to specific diseases and drawing on existing frameworks such as Ragas and CMB.
Building an intelligent medical history collection system using voice interaction, natural language processing, and large model technologies. Speech recognition and NLP techniques collect patient information through multi-round dialogue with strategies including questioning, paraphrasing, feeling reflection, and self-disclosure. Physicians supervise and validate the model through Reinforcement Learning from Human Feedback to improve collection efficiency. After collection, the model extracts key information (chief complaint, history of present illness, past medical history, personal history, allergy history) and uses Retrieval-Augmented Generation (RAG) combined with guidelines for preliminary triage.
Building an intelligent clinical auxiliary diagnosis system with multi-source data fusion integrating heterogeneous medical records, laboratory reports, and imaging text reports using OCR and natural language processing (NLP). Using chain-of-thought and reasoning capabilities of large language models combined with knowledge graph retrieval to generate evidence-based diagnostic recommendations. Physicians provide feedback through a data annotation interface to fine-tune the system, with customization settings to meet individual physician needs.
Building a medical record generation system enabling automatic entry of patient information and automatic medical record generation using large language models for entity recognition, text classification, and semantic understanding. The system generates records compliant with writing standards including chief complaint, history of present illness, past medical history, personal history, marital and reproductive history, family history, and auxiliary examinations. A quality control system automatically detects generated records and, when core fields are missing, alerts physicians through highlighting and submission blocking, prompting them to make necessary additions and modifications.
Providing a solution for consultation audio data collection and processing to support intelligent history collection, rapid triage, and clinical auxiliary diagnosis. The collected audio is converted to accurate textual content with preliminary analysis to identify key medical information (symptoms, medications, treatment processes). The study addresses speech recognition challenges including specialized terminology, dialects, accents, and elderly speech characteristics, using refined speech processing algorithms and deep learning approaches. Strict data protection measures (encryption, anonymization) are applied throughout collection, transmission, storage, and processing to ensure patient data security.
Validating the generalizability of the five intelligent diagnosis and treatment systems (gastric cancer, chronic gastritis, GERD, coronary artery disease, stroke) by transferring the fine-tuned systems from the training center to other centers for external cross-validation. Feasibility and accuracy are evaluated by comparing information collected from large model-patient interactions and physician-patient interactions across different center environments.
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
The medical large language model fine-tuning technology platform is established based on the ChatGLM foundation model. The overall research includes medical big data construction, medical knowledge graph construction, medical large model training and fine-tuning, and large model application platform development and deployment.
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
Development of a high-quality doctor-patient dialogue audio collection and recognition system for complex medical scenarios, providing the data foundation for medical history collection and medical record generation. The system uses a self-developed multi-microphone array and sound source localization technology for precise speaker positioning, with wiener filtering and echo cancellation to reduce environmental noise. Deep learning is used for speech recognition, and large models are applied for post-processing to correct recognition errors.
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
This study has accumulated multimodal medical data including electronic medical records, medical images, laboratory results, and genetic information, as well as medical guidelines, expert consensus, clinical databases, medical literature, encyclopedias, patents, and doctor-patient dialogue data. A multi-source heterogeneous database is constructed covering images, audio, video, text, and tables from different stages inside and outside the hospital. A self-developed text recognition tool digitizes PDF guidelines and literature while preserving tables, images, and LaTeX-format formulas. A medical machine translation tool translates English literature into Chinese. Through knowledge extraction, fusion, and completion, a tens-of-millions-level medical knowledge graph is built. Data cleaning, standardization, transformation, annotation, and augmentation convert raw data into forms suitable for model training and diverse clinical applications
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
Based on the ChatGLM model, this study develops a full-cycle method and platform for domain-specific large model fine-tuning using technologies including Continue Pretrain, Multi-task Instruction Pretrain, Reinforcement Learning, In-context Learning, Mixture-of-Experts, Chain-of-Thought, Parameter-efficient fine-tuning, and prompt-tuning, establishing a comprehensive, objective, and effective medical model fine-tuning system in combination with real-world application scenarios. The technology can be applied to large model fine-tuning across different diseases, meeting the needs of pre-consultation, in-consultation, and post-consultation scenarios. Combined with task-specific annotation data, instruction fine-tuning enhances capabilities in knowledge question answering, summarization, knowledge extraction, medical record generation, and auxiliary diagnosis. The model is deployed and further optimized through reinforcement learning with human feedback from physicians and patients
Time frame: From enrollment to completion of the knowledge graph construction, assessed up to 2 years
Prompt Learning and Few-shot Learning enable few-shot knowledge extraction using human-in-the-loop iteration (pre-annotation, manual verification, continuous training). Self-developed knowledge linking and alignment algorithms integrate knowledge from guidelines, textbooks, consensus, and medical records. A knowledge graph extraction method combining local semantic and global structural information is proposed for digestive, cardiovascular, and cerebrovascular diseases. Entity recognition and representation are performed on the large model, and graph neural networks learn from existing knowledge graphs to mine entity relationships. Under an active learning framework, high-value data are manually annotated, enabling rapid model iteration with low cost to construct disease-specific knowledge graphs for auxiliary diagnosis and other scenarios
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
Development of a medical history collection system based on the medical large model to achieve rapid and accurate history collection through patient dialogue. The system allows patients to upload medical records and examination results, extracts key information through recognition and extraction, and displays it to physicians during consultations to improve work efficiency.
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
Development of a diagnostic support system for corresponding diseases based on the disease-specific large model to address the challenges of low medical homogenization and high diagnostic inconsistency. The system incorporates both the diagnostic knowledge from disease-specific guidelines and the clinical experience of specialists from tertiary hospitals. Through this system, primary care hospitals are empowered to improve the overall level of medical services.
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
Development of a medical record generation and quality control system based on pre-consultation, history collection, and face-to-face consultation data. The system directly generates medical records from doctor-patient dialogue data using the medical large model, performs automatic quality inspection and prompting, and generates high-quality inpatient records. The system integrates with hospital systems to reduce the burden on physicians in medical record writing. Additionally, the system can perform quality inspection and automatic optimization of historical medical records to improve overall hospital data management capabilities.
Time frame: From enrollment to completion of the cross-center clinical validation, assessed up to 3 years
Validation of the generalizability of the five intelligent diagnosis and treatment systems (gastric cancer, chronic gastritis, gastro esophageal reflux, coronary artery disease, stroke) by transferring the fine-tuned systems from the training center to other centers for external cross-validation. Feasibility and accuracy are evaluated by comparing information collected from large model-patient interactions and physician-patient interactions across different center environments.
Time frame: From enrollment to completion of the platform development, assessed up to 2 years
On the basis of the disease-specific large model, the entire process of disease diagnosis and treatment is empowered, including: (1) automatic collection of patient disease information; (2) the diagnostic support system generates precise diagnostic recommendations based on patient information to assist physicians in improving consultation efficiency; and (3) improvement of medical record standardization.
Contact information is provided by the study sponsor or research team.
Beijing Friendship Hospital
Other
Development and Validation of Fine-Tuning Techniques for Large Models in Gastric, Cardiovascular and Cerebrovascular Diseases
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT03847753
Acquired Immunodeficiency Syndrome, Allergy
View Trial DetailsNCT04697186
Chronic Gastritis, Digestive System Diseases
Xi'an, Shaanxi, China
View Trial DetailsNCT03609892
Chronic Gastritis, Digestive System Diseases
Xi'an, Shaanxi, China
View Trial DetailsNCT02219529
Chronic Gastritis, Digestive System Diseases
Shanghai, China
View Trial Details