The First Affiliated Hospital of Wenzhou Medical University
Wenzhou, Zhejiang, China
Location status: Recruiting
Location contact
Qi Lin
CONTACT
Yikai Chen
CONTACT
NCT Number: NCT07707232
This observational study will develop and validate a large language model-assisted workflow for imaging cTNM staging annotation and uncertainty recognition in prostate cancer using Chinese PSMA PET/CT report texts generated during routine clinical care. The study will use de-identified report texts and necessary baseline clinical information only. No additional imaging examination, blood test, treatment, or follow-up visit will be assigned for this study.
The main objective is to evaluate whether a locally or institutionally controlled large language model can identify report-derived imaging cT, cN, and cM categories, extract supporting evidence from the original report, and recognize uncertainty expressions. Model performance will be assessed using an internal independent validation set, external validation reports from two collaborating hospitals, and a prospective validation set of 100 consecutive routine PSMA PET/CT reports. A human-AI comparison will also be performed using physicians from urology and imaging-related specialties with different seniority levels.
Interested in participating?
Request Info18 year and older
Male
Observational
Wenzhou, Zhejiang, China
Location status: Recruiting
Qi Lin
CONTACT
Yikai Chen
CONTACT
This is a multicenter observational diagnostic accuracy validation study based on Chinese PSMA PET/CT report texts from patients with prostate cancer or suspected prostate cancer. The study is not designed to evaluate a drug, device, surgical procedure, or imaging intervention. PSMA PET/CT examinations will be performed as part of routine clinical care, and the study will only analyze de-identified report texts and necessary baseline information after the reports have been finalized.
The study consists of retrospective and prospective components. Retrospectively, approximately 4,000 PSMA PET/CT reports from the First Affiliated Hospital of Wenzhou Medical University will be systematically annotated to construct a research database. An internal independent validation set of 300 reports, not used for model development or prompt optimization, will be used to evaluate the performance of the large language model. The reference standard for this 300-report validation set will be established by two experienced urologists through joint annotation, with adjudication by a nuclear medicine expert when needed. External validation will be performed using 110 de-identified reports from the First Affiliated Hospital of Ningbo University and 102 de-identified reports from Liuzhou People's Hospital. In addition, after ethics approval, 100 consecutive routine PSMA PET/CT reports from the First Affiliated Hospital of Wenzhou Medical University will be prospectively included to evaluate the accuracy and operational stability of the frozen model and prompt versions.
The large language model workflow will be deployed locally or in an institutionally controlled environment. The model will be instructed to generate structured JSON outputs, including cT_report, cN_report, cM_report, cT_uncertain, cN_uncertain, cM_uncertain, evidence_T, evidence_N, and evidence_M. The target task is report-derived imaging cTNM staging annotation, not pathological TNM staging or overall AJCC stage grouping. The model output will be used only for research evaluation and methodological analysis and will not be used for clinical diagnosis, treatment decision-making, or patient notification.
A human-AI comparison will be conducted on the 300-report internal validation set. Eight human evaluators from urology and imaging-related specialties, including trainees, residents, attending physicians, and associate chief physicians, will independently annotate the reports before and after learning the annotation manual. Annotation time will be recorded for each round. The performance of human evaluators and the large language model will be compared against the expert consensus reference standard.
The main outcome will be the accuracy of the large language model in identifying cT, cN, and cM categories from Chinese PSMA PET/CT reports. Secondary outcomes will include precision, recall, F1-score, macro-F1, micro-F1, complete cTNM triplet matching rate, uncertainty recognition performance, evidence extraction quality, human-AI comparison results, annotation time, external validation performance, prospective validation performance, and error type distribution. Error analysis will focus on local tumor extent, regional versus non-regional lymph node boundaries, bone and visceral metastasis recognition, equivocal wording, treatment-related context, benign or inflammatory alternatives, and lesions not attributable to prostate cancer.
Healthy volunteers accepted: No
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
A locally or institutionally controlled large language model workflow will analyze de-identified Chinese PSMA PET/CT report texts and generate structured outputs for report-derived imaging cTNM staging annotation, uncertainty recognition, and supporting evidence extraction. This workflow is used only for research evaluation and methodological analysis. It will not assign any examination, treatment, medication, procedure, or follow-up to participants, and it will not guide clinical diagnosis or treatment decisions.
Time frame: After freezing the model and prompt versions, through completion of internal, external, and prospective validation, up to 18 months
The primary outcome is the accuracy of the large language model in identifying report-derived imaging cT, cN, and cM categories from de-identified Chinese PSMA PET/CT report texts. The LLM-generated cT_report, cN_report, and cM_report will be compared with the expert consensus reference standard. Accuracy, precision, recall, F1-score, macro-F1, micro-F1, complete cTNM triplet matching rate, and confusion matrices will be calculated in the internal 300-report validation set, external validation sets, and prospective 100-report validation set.
Time frame: After freezing the model and prompt versions, through completion of all validation analyses, up to 18 months.
This outcome measures the component-level accuracy of the large language model in recognizing uncertainty labels for report-derived imaging cTNM staging. The LLM-generated cT_uncertain, cN_uncertain, and cM_uncertain labels will be compared with the expert consensus reference standard. Accuracy will be calculated as the number of correctly classified uncertainty labels divided by the total number of component-level uncertainty labels across all reports. The three uncertainty components will be aggregated into one percentage value.
中文对应
Time frame: During pre-training and post-training human annotation rounds and LLM batch inference, up to 18 months.
This outcome measures the proportion of reports for which the complete report-derived imaging cTNM triplet assigned by human evaluators and by the large language model exactly matches the expert consensus reference standard. A report will be counted as correct only when all three components, cT_report, cN_report, and cM_report, are correct. The result will be reported as the percentage of reports with complete cTNM triplet agreement. Results will be summarized separately for pre-training human annotation, post-training human annotation, and LLM batch inference.
Time frame: During pre-training and post-training human annotation rounds and LLM batch inference, up to 18 months.
This outcome measures the mean time required to complete report-derived imaging cTNM annotation per report. For human evaluators, annotation time will be recorded during the annotation rounds and divided by the number of annotated reports. For the LLM workflow, batch inference time will be divided by the number of processed reports. Results will be reported separately for pre-training human annotation, post-training human annotation, and LLM batch inference.
Contact information is provided by the study sponsor or research team.
Qi Lin
CONTACT
Yikai Chen
CONTACT
First Affiliated Hospital of Wenzhou Medical University
Other
Large Language Model-Assisted Imaging cTNM Staging Annotation and Uncertainty Recognition for Prostate Cancer Based on Chinese PSMA PET/CT Reports
Acronym: PSMA-LLM-cTNM
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT05751434
Genital Diseases, Genital Diseases, Male
Los Angeles, California, United States
View Trial DetailsNCT07028853
Cancer of the Prostate, Genital Diseases
Huntsville, Alabama, United States
View Trial DetailsNCT05765500
Genital Diseases, Genital Diseases, Male
Boston, Massachusetts, United States
View Trial DetailsNCT07203482
Genital Diseases, Genital Diseases, Male
Brooklyn, New York, United States
View Trial Details