📊 Breast cancer AI performance varied by data type
📊 Breast cancer AI performance varied by data type
In a study comparing AI approaches for breast cancer detection across structured clinical datasets and imaging datasets, performance varied sharply by data type: machine learning models reached up to 98% accuracy on structured data, while deep learning models led image tasks with 97.88% accuracy on histopathology and 92.00% on ultrasound. The study also found calibration differed by model family, with Logistic Regression, Gradient Boosting, and Random Forest among the best-calibrated structured-data models, while transfer-learning deep networks were better calibrated on histopathology images.
Why It Matters To Your Practice
AI may be useful as a complementary diagnostic support tool when breast imaging is inconclusive or reader variability is a concern.
The findings suggest one model type is no longer enough for breast cancer AI across workflows; performance depends heavily on whether the input is clinical tabular data or images.
Calibration matters clinically because well-calibrated probability estimates may better support risk discussions and downstream decisions.
Clinical Implications
For structured clinical data, ensemble and probabilistic machine learning models performed strongly, with accuracy up to 98%.
For imaging tasks, deep learning models outperformed traditional machine learning, especially on histopathology images and ultrasound.
Logistic Regression, Gradient Boosting, and Random Forest showed stronger calibration than Decision Tree and Naive Bayes on structured datasets.
Transfer-learning architectures appeared better calibrated than a from-scratch CNN baseline and some models tested on the smaller BUSI ultrasound dataset.
Insights
The analysis covered multiple structured datasets, including the Wisconsin Breast Cancer Dataset, CSAW-CC, and SEER, plus mammography, histopathology, and ultrasound image datasets.
Deep learning models assessed included CNN, ResNet50, VGG16, and DenseNet.
All reported performance metrics included 95% confidence intervals, and calibration was assessed with Expected Calibration Error and reliability diagrams.
Smaller datasets appeared more vulnerable to weaker calibration and less consistent model performance.
The Bottom Line
AI for breast cancer detection is promising, but clinicians should match the model approach to the data source rather than expect one system to generalize across all settings.
Prospective validation across diverse patient populations is still needed before broad clinical adoption.