Saeide Danaei, Zahra Dehghanian, Elahe Meftah, Nariman Naderi, Seyed Amir Ahmad Safavi-Naini, Faeze Khorasanizade, Hamid R. Rabiee
5 min
Artificial intelligence (AI) has become a cornerstone of modern medical imaging, particularly for abdominal CT analysis. However, the efficacy of these models is fundamentally tied to the quality, diversity, and representativeness of the training data. This systematic review evaluates 46 publicly available abdominal CT datasets, comprising over 50,000 studies, to determine their suitability for real-world clinical deployment.
The review highlights two primary issues: high levels of data redundancy and significant geographic skew. Approximately 59% of the studies examined were found to be reused across multiple datasets, which increases the risk of data leakage and overfitting. Furthermore, there is a pronounced Western bias, with 75.3% of datasets originating from North America and Europe. This lack of global representation, particularly the total absence of data from many parts of Africa, South Asia, and the Middle East, suggests that models trained on these resources may fail to generalize effectively in diverse, resource-limited healthcare environments.
When evaluating the 19 datasets containing at least 100 cases, the researchers identified significant risks related to domain shift (63%) and selection bias (57%). These biases indicate that models often struggle when applied to data from different clinical settings or patient populations. Additionally, the datasets show a heavy focus on tumor-centric tasks, often neglecting common non-neoplastic conditions. This narrow focus limits the utility of current AI models in routine diagnostic scenarios where a broader range of pathologies is encountered.
To improve the clinical robustness of AI in abdominal imaging, the authors advocate for multi-institutional collaboration and the adoption of standardized protocols. They emphasize that future dataset curation must prioritize the inclusion of diverse patient populations and older-generation imaging technologies, which are more representative of the infrastructure found in many global hospitals. By addressing these systemic biases, the research community can move toward creating more equitable and reliable AI tools for global healthcare.
This review systematically searched publicly available abdominal CT datasets and critically evaluates suitability for artificial intelligence (AI) applications in clinical settings. We examined 45 publicly available abdominal CT datasets (47,049 studies). Across all 45 datasets, we found substantial redundancy (51% case reuse) and a Western/geographic skew (75.3% from North America and Europe). A bias assessment was performed on the 22 datasets with more than 100 cases; within this subset, the most prevalent high-risk categories were racial bias (with a score of 16 out of 22) and selection bias (with a score of 15 out of 22), both of which may undermine model generalizability across diverse healthcare environments-particularly in resource-limited settings. To address these challenges, we propose targeted strategies for dataset improvement, including multi-institutional collaboration, adoption of standardized protocols, and deliberate inclusion of diverse patient populations and imaging technologies. These efforts are crucial in supporting the development of more equitable and clinically robust AI models for abdominal imaging.
Alex: [nodding] Precisely. The model treats image quality differences as features rather than nuisance variables. Because the training data lacks technical heterogeneity, the model cannot generalize across the hardware landscape that actually exists in global clinical settings.
Sam: [probing] And the selection bias piece — if these datasets are skewed toward patients already flagged for surgery or advanced therapy, are the models essentially only learning to recognize the most obvious, late-stage presentations?
Alex: [deliberate] That is exactly the implication. If your training set is composed primarily of patients who have already cleared multiple clinical filters, the model never encounters the subtle, early-stage pathologies that are arguably the most important to catch. It overfits to the conspicuous cases and fails quietly on the ambiguous ones. [[RP_SECTION:implications-for-model-generalizability|Implications for Model Generalizability]]
Sam: [thoughtful] So the "nutrition label" framing the authors propose — that is really about transparency in the patient population as much as the imaging hardware. A researcher who does not know the disease spectrum in their training set is essentially flying blind on generalizability.
Alex: [measured] That is the core of the problem. And it compounds: because so many of these datasets are recycled, a bias introduced in one foundational dataset propagates through the literature. A model trained on a biased base dataset, fine-tuned on another dataset that partially overlaps with it, evaluated on a benchmark that draws from the same pool — the apparent performance gains are not evidence of genuine robustness.
Sam: [summarizing] So the state-of-the-art metrics we see in these papers are partly a measure of how well a model memorizes a narrow, Western-centric slice of clinical reality — and the benchmark is too close to the training distribution to reveal that.
Alex: [slower, for emphasis] That is a fair characterization of what the evidence supports. The authors argue that without multi-institutional collaboration and deliberate inclusion of diverse imaging technologies and patient populations, we risk building AI that exacerbates health disparities rather than addressing them. [[RP_SECTION:future-of-data-rigor|Future of Data Rigor]]
Sam: [reflective] Which means the next step for the field is not simply acquiring more data — it is better-indexed, more representative data, with provenance treated as a first-class variable.
Alex: [concluding, quiet conviction] Exactly. The authors' core recommendation is that we apply the same rigor to dataset provenance that we currently apply to model architecture. Right now, the field is meticulous about ablating architectural choices and reporting confidence intervals on benchmark performance — but the data pipeline that feeds those benchmarks is largely taken on faith. That asymmetry is where the fragility lives. Thanks for listening to ResearchPod.