K² · Artificial intelligence
Cancer-detecting AI misses tumors in young Black women it rarely trained on
The authors built BenchX, a large-scale open benchmark of 85,355 CT scans to evaluate AI models for cancer detection and localization.
QC
WL
Qi Chen, Wenxuan Li, Pedro R. A. S. Bassi et al.
17 authors · cs.CV
The 30-second scan
An AI that spots tumors on scans can look excellent on paper and still fail the exact patients it sees least.
- The benchmark systematically evaluates 12 tumor-detection AI models across tumor size, location, patient subgroup, and imaging protocol.
- The authors leverage large language models (LLMs) to extract and organize subgroup information from clinical data, making the analysis scalable and reproducible.
- The benchmark reveals that current state-of-the-art AI models, optimized for average accuracy, perform poorly in rare or underrepresented subgroups, such as young, female African Americans.
85,355 CT scans in the benchmark12 tumor-detection AI models evaluated