Design. LTC-VALID is a prospective, two-center, single-arm, two-phase clinical validation study of a CE-marked class IIa medical device software intended for the automated interpretation of laboratory test results.
Phase structure. Phase I (609 participants, months 1 to 9) is exploratory and serves algorithm development and gap identification using version 1.0 of the software. An interposed optimization stage (months 9 to 10) produces a frozen version 2.0. Phase II (290 participants, months 11 to 13) is confirmatory. All outcome measures listed below are collected identically in both phases and are reported separately by phase. Results from the two phases are not pooled, because the two phases evaluate different versions of the software. The pre-specified confirmatory hypothesis test for the co-primary measures is applied to Phase II data; Phase I results for the same measures are exploratory and are reported descriptively.
Procedures. Eligible participants provide written informed consent, are referred for a mandatory basic laboratory panel and one or two of 29 specialist laboratory panels, and provide a single blood draw at a certified laboratory collection point. After the results become available, the participant completes a dynamically generated electronic medical history questionnaire. The software produces one interpretation per laboratory result; interpretations are not aggregated by the software. The attending physician records an independent clinical assessment of the same data while blinded to the software output; this assessment is locked before the interpretation report is released to the participant. Participants complete a questionnaire evaluating the report, and physicians complete a form evaluating the completeness and relevance of the automated medical history.
Reference standard and comparators. An independent expert physician receives the complete documentation and establishes an own reference assessment before reviewing the assessments to be compared. The assessments of the software, of the attending physician and of large language models are presented in random order and blinded as to authorship. The comparators are comparators of assessment, not study arms; the study is single-arm and no randomization or control group is used.
Reporting standard. The primary analysis follows the Standards for Reporting of Diagnostic Accuracy Studies (STARD). The study is a diagnostic accuracy study and not a study of clinical effectiveness.
Statistical approach. The two primary outcome measures are co-primary and are evaluated using an intersection-union test; both must meet their pre-specified criteria. The pre-specified criteria apply to Phase II: for the safety measure, the lower bound of the one-sided 95 percent confidence interval is at or above 95 percent; for the accuracy measure, at or above 90 percent. Confidence intervals for proportions are calculated using exact methods (Clopper-Pearson). Phase I is exploratory and its data are reported descriptively. Sensitivity for the rare urgency categories is reported with confidence intervals and event counts as a secondary, non-confirmatory measure, because its denominator is not sufficient for formal hypothesis testing at the planned sample size.