Background: Artificial intelligence (AI) models are increasingly evaluated for diagnosis in medical imaging and pathology, but reported performance depends on validation strategy. Methods: PubMed, Embase, Cochrane Library, IEEE Xplore and Web of Science were searched for open-access English-language studies from January 2021 to February 2026 reporting the area under the receiver operating characteristic curve (AUC) of a diagnostic AI model. AUC values were summarized descriptively by validation strategy, architecture and modality. Studies with 2×2 data from an external validation set were pooled with a bivariate random-effects model. Risk of bias was assessed with the Quality Assessment of Diagnostic Accuracy Studies-2 tool. Results: Of 1,503 records screened, 170 studies (one model each) were included. AUC was ≥0.90 in 119 studies (70.0%); mean AUC was 0.945 (internal and external validation), 0.929 (internal), 0.912 (external) and 0.798 (four randomized trials). Seven externally validated studies of unrelated conditions provided 2×2 data (summary sensitivity 92.1% (95% confidence interval 88.5–94.6), specificity 90.1% (81.3–95.0); wide prediction region). Risk of bias was frequently unclear (patient selection 71.8%, flow and timing 64.7%). Conclusions: High reported AUC values reflect discrimination in selected, mostly retrospective datasets and do not establish clinical effectiveness or readiness for routine use.