Hopital Européen Georges Pompidou
Paris, France
Location status: Recruiting
NCT Number: NCT07739121
BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.
Interested in participating?
Request Info18 year and older
All sexes
Observational
Paris, France
Location status: Recruiting
BEACON is a prospective, multicentre, blinded benchmark using automated, criteria-based scoring. It is built on three design decisions that distinguish it from the existing literature: (i) synthetic, standardised cases remove the record-completeness variability that confounds retrospective comparisons and allow the identical input to be given to every board and every model; (ii) two independent tumour boards per localisation let human-human agreement be measured rather than assumed; and (iii) a guideline matrix, locked a priori, provides an objective anchor applied identically to human and model recommendations.
Reference standard. For each case-domain, a guideline matrix (guideline-recommended / acceptable / unsupported options per case-domain; ESMO, NCCN), locked and time-stamped before data collection, is applied identically to boards and models.
Five decision domains. Every recommendation is decomposed into D1 Intent, D2 Surgery, D3 Radiotherapy, D4 Systemic therapy (class + line), and D5 Work-up & biomarkers before any comparison.
Healthy volunteers accepted: No
Only the study team can determine whether someone qualifies for participation.
Inclusion criteria
Exclusion criteria
Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.
Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.
Time frame: Assessed once at central scoring, after data collection (~October 2026)
For each recommendation domain and each LLM, proportion of LLM recommendation concordant with locked guidelines
Time frame: Up to October 2026
Time frame: Up to October 2026
Each recommendation domain, decomposed into the five decision domains and scored per domain on an ordinal scale (2 = complete concordance; 1 = partial concordance; 0 = discordance).
Time frame: Up to October 2026
Agreement between the two independent tumour boards scored per recommendation domain
Time frame: Up to October 2026
Proportion of case-domains where the two tumour boards give different categorical recommendations
Time frame: Up to October 2026
Proportion of required domains addressed (LLM and tumour boards)
Time frame: Up to October 2026
Proportion of critical omissions (LLM and tumour boards)
Time frame: Up to October 2026
Proportion of recommendation corresponding to over- or under-treatment
Contact information is provided by the study sponsor or research team.
Jean-Emmanuel Bibault, MD PhD
CONTACT
01 56 09 34 06 ext. +33
Jérôme Lambert, MD PhD
CONTACT
0142499742 ext. +33
Assistance Publique - Hôpitaux de Paris
Other
Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning
Acronym: BEACON
OpenTrials presents study information sourced from ClinicalTrials.gov. The official registry record should be consulted for the latest information.
View the official ClinicalTrials.gov record (opens in a new tab)This listing is for discovery and informational purposes only. It is not medical advice, does not guarantee that a study is recruiting, and does not determine eligibility. Contact the study team and a qualified healthcare professional when considering participation.
Published trials that share one or more normalized conditions with this study.
NCT02897778
Breast Diseases, Breast Neoplasms
San Antonio, Texas, United States
View Trial DetailsNCT00120939
Adnexal Diseases, Astrocytoma
Pittsburgh, Pennsylvania, United States
View Trial DetailsNCT06379880
Breast Diseases, Breast Neoplasms
Vandœuvre-lès-Nancy, De, France
View Trial DetailsNCT00978250
Breast Diseases, Breast Neoplasms
Davis, California, United States
View Trial Details