The Alzheimer's Disease Neuroimaging Initiative (ADNI) is widely used to train machine learning models for Alzheimer's disease, yet whether it provides consistent ground truth for predictive modeling has not been systematically tested. In this paper, we showed that ADNI contains three interacting sources of bias with direct implications for machine learning: (a) diagnostic label inconsistency, (b) technical measurement drift, and (c) longitudinal survivor bias. A substantial proportion of cases, particularly within intermediate stages, fall outside ADNI's own diagnostic thresholds. MRI field strength and evolving processing pipelines introduce significant technical variability in hippocampal volume, while cohort survivor bias arising from differential retention of participants across phases further distorts longitudinal estimates of disease progression. These findings indicated that ADNI does not provide the stable, internally consistent labels often required in machine learning applications. We proposed a practical framework for diagnostic validation, feature harmonization, and cohort accounting, offering guidance for building more robust and biologically meaningful predictive models from large-scale neuroimaging cohorts.
Smith, P., Richardson, H., Robertson, K., Dibble, A. J., Dalby, C., Svanera, M.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 5
- Comments 0
