RecruitingNCT07739121

Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations

Benchmarking AI for Clinical Oncology decisioNmaking (BEACON): A Prospective, Multicentre, Blinded Evaluation of Frontier Large Language Models Against Multidisciplinary Tumour Board Recommendations in Oncology Treatment Planning


Sponsor

Assistance Publique - Hôpitaux de Paris

Enrollment

100 participants

Start Date

May 1, 2026

Study Type

OBSERVATIONAL

Conditions

Summary

BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.


Eligibility

Min Age: 18 Years

Plain Language Summary

Simplified for easier understanding

This clinical trial is studying Frontier large language models and Multidisciplinary tumour boards for people with artifical intelligence, breast neoplasms, and other related conditions. The study is currently recruiting participants at 1 location. People eligible for this study include aged 18 Years and older.

This summary was AI-generated to explain the trial in plain language. It is not medical advice. Always discuss eligibility with your doctor before enrolling in a clinical trial.

Interested in this trial?

Get notified about updates and connect with the research team.

Interventions

OTHERMultidisciplinary tumour boards

Two independent tumour boards per localisation (10 boards in total) issue a categorical recommendation for every synthetic case. Where both boards agree, their consensus defines the reference standard; where they differ, the case-domain is classified as EQUIPOISE and analysed separately.

OTHERFrontier large language models

Five frontier LLMs (GPT-5.6, Claude Fable 5, Gemini 3.1 Pro, DeepSeek V4 Pro, Llama 4 Maverick) each receive the identical structured input for every case, three times in independent sessions, under locked prompts, versions and settings.


Locations(1)

Hopital Européen Georges Pompidou

Paris, France

View Full Details on ClinicalTrials.gov

For the most up-to-date information, visit the official listing.

Visit

NCT07739121


Related Trials