Screening works, but it strains the people doing it.
Breast cancer is the most commonly diagnosed cancer among women worldwide, and early detection is the strongest lever on outcomes. Yet the standard radiodiagnostic toolkit of mammography, ultrasound, and MRI carries real limitations: variable sensitivity, false positives and false negatives, and reduced accuracy in dense breast tissue, where lesions are hardest to see.
On top of that sits a human factor. Large screening programs generate image volumes that place a heavy cognitive load on radiologists, and interpretation varies between readers. The review asks whether AI, particularly deep learning, can improve diagnostic accuracy and efficiency, and examines the practical barriers, from data quality to ethics and regulation, that shape whether it should be deployed.
A systematic search, filtered for real comparisons.
The review searched CINAHL, Cochrane Library, Medline, Health and Psychological Instruments, and Health Source for peer-reviewed studies published between 2020 and 2024, combining terms such as artificial intelligence, breast cancer detection, radiology, deep learning, computer-aided detection, and diagnostic accuracy with Boolean operators.
Studies qualified only if they were peer reviewed, available in full text in English, and compared AI against traditional diagnostic methods. Papers that only described model mechanics without a clinical comparison, or that centered on pathology rather than radiology-based detection, were excluded. After title, abstract, and full-text screening, 31 core articles were selected and organized in a literature synthesis matrix across themes: diagnostic accuracy, AI versus traditional methods, multi-modal integration, and implementation challenges.
Accuracy gains are real, and so are the trade-offs.
Diagnostic accuracy
Across modalities, AI models posted strong performance on detection tasks. A Random Forest classifier reached 98.83% accuracy classifying mammograms as benign or malignant (Ammar et al., 2024). A multicenter CNN detected invasive carcinoma in breast whole slide images with up to 97.6% accuracy, holding up across clinical centers through transfer learning (Peyret et al., 2023). In digital breast tomosynthesis, AI reached sensitivity up to 95.7% in a multi-institutional study (Konz et al., 2023), and AI analysis of contrast-enhanced mammography differentiated benign from malignant lesions with 95.83% accuracy (Antonella et al., 2022). In double-reading programs, the GMIC and GLAM models identified cancers human readers had missed, at sensitivities of 82.98% and 79.79% (Jiang et al., 2024).
AI versus radiologists: a specificity story
Head-to-head data show the gains are not uniform across metrics. In a screening cohort comparison, standalone AI was roughly comparable to radiologists on sensitivity but clearly stronger on specificity and recall behavior, which is where much of the patient and system burden lives (Mi-ri et al., 2024).
| Metric | AI | Radiologists |
|---|---|---|
| Sensitivity (%) | 67.1 | 69.9 |
| Specificity (%) | 93.0 | 77.6 |
| Positive predictive value (%) | 1.5 | 0.5 |
| Recall rate (%) | 7.1 | 22.5 |
| Area under the curve | 0.80 | 0.74 |
The balancing evidence matters too. In the BreastScreen NSW program, AI matched or exceeded radiologists on sensitivity but showed lower specificity in that workflow, increasing recalls and arbitration reads (Warner-Smith et al., 2024). Whether AI reduces or adds burden depends on the pathway it is deployed in, not just the model.
Workload and screening efficiency
Efficiency findings were among the most consistent. A meta-analysis of AI triage found radiologist workload reductions of up to 68% while preserving 93.1% sensitivity (Xavier et al., 2024). A population screening simulation compared deployment pathways directly: the AI band-pass configuration cut false positives by 13.3% and human reads by 80.7% relative to the current reader system (Helen et al., 2024).
| Scenario | True positives | False negatives | False positive reduction (%) | Human read reduction (%) |
|---|---|---|---|---|
| Current reader system | 1,061 | 268 | 0 | 0 |
| AI reader-replacement | 1,094 | 235 | 6 | 48 |
| AI band-pass | 1,086 | 243 | 13.3 | 80.7 |
| AI triage | 1,026 | 303 | -7.3 | 46 |
Prospective evidence points the same direction: an AI-assisted mammography tool detected an additional 0.7 to 1.6 cancers per 1,000 individuals screened, skewing toward smaller and invasive cancers, though with increased recall rates that need workflow tuning (Kontos, 2024). Beyond radiology, an AI assist raised pathologists' sensitivity for lymph node metastases from 74.5% to 93.5% while cutting reading time by 55% (Retamero et al., 2024), and AI-guided selection made supplemental MRI screening more cost-effective by triaging high-risk patients (Salim et al., 2024).
Multi-modal integration and advanced techniques
The most advanced systems combine mammography with ultrasound, MRI, genomics, and patient history. Reviewed studies showed ultrasound-based AI predicting molecular markers such as HER2 and Ki67 (Fu et al., 2024), MRI radiomics estimating recurrence risk in ER+/HER2- patients as a non-invasive complement to assays like Oncotype Dx (Chiacchiaretta et al., 2023), and transfer learning with radiomics classifying breast tumors at a balanced accuracy of 0.964 with an AUC of 0.981 (Tian et al., 2024). Hybrid feature-selection models reached 98.25% accuracy for initial diagnosis and 93.27% for recurrence prediction (Pati et al., 2024). The direction of travel is holistic: models that read more than one signal support earlier and more personalized decisions.
What the evidence does not yet establish.
- Mostly retrospective designs. Most models were validated on pre-collected datasets; prospective trials in live screening conditions remain the critical missing evidence.
- Training data bias and generalizability. Datasets often underrepresent demographic diversity in breast density, ethnicity, and risk profile, which risks unequal performance across populations. Federated learning and multi-center, multi-vendor training data are the leading remedies.
- Black-box interpretability. Limited insight into how models reach conclusions undermines clinician trust; explainable AI methods are needed for adoption.
- Specificity trade-offs in some workflows. Depending on deployment pathway, AI can raise recall and arbitration volumes, adding burden instead of removing it.
- Regulation and trust. Validation and reporting frameworks such as the EU trustworthy AI guidelines, TRIPOD-AI, and CONSORT-AI need to be embedded in policy. Patient acceptance is conditional: 64% of women supported AI-assisted reads, but only alongside human oversight and transparency (Holen et al., 2024).
A second reader, not a replacement.
The review's throughline is that AI performs best as a decision-support layer: a concurrent or second reader that screens out low-risk cases, flags subtle abnormalities, and frees radiologists to spend judgment where it matters. Radiologists supply the contextual understanding, clinical reasoning, and patient history that current models lack.
For health informatics practice and policy, the implications are concrete: standardized diagnostic criteria can reduce interpretive variability, AI-specific validation standards belong in national screening policy, clinician training on AI belongs in medical education, and datasets must be representative before deployment can be equitable. Staged implementations with human oversight, transparency, and opt-out options are the pattern patients and clinicians say they will trust.
“The evidence suggests that AI can play a transformative role in breast cancer diagnostics, but it is most effective when used as a decision-support tool rather than as a replacement for human expertise.”
Patel, R. (2024). Exploring the role of artificial intelligence (AI) in enhancing the accuracy and efficiency of breast cancer detection and diagnosis. Systematic review, M.S. Healthcare Informatics capstone, Sacred Heart University. Recognized by New England HIMSS, 2025.
This brief condenses the full paper, which reviews 31 core peer-reviewed studies with complete methods, synthesis, and references. The manuscript is being prepared for peer-reviewed journal submission.