Department of Planning, Research and Statistics, Federal Neuropsychiatric Hospital, Kano, Nigeria
© 2026, Korean Society of Epidemiology
This is an open-access article distributed under the terms of the Creative Commons Attribution License (http://creativecommons.org/licenses/by/4.0/), which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
Conflict of interest
The author has no conflicts of interest to declare for this study.
Funding
None.
Acknowledgements
None.
Author contributions
All work was done by Musa SA.
| Aspects | Phase I: Classical probability-based sampling | Phase II: Digitally mediated sampling | Phase III: AI-enhanced adaptive sampling | Hybrid frameworks (recommended future) |
|---|---|---|---|---|
| Core emphasis | Design-based inference and randomization | Scalability and large-scale data access | Dynamic, information-driven selection | Integration of probabilistic foundations with adaptive algorithms |
| Key methods | Simple random, stratified, cluster, and multistage sampling; non-probability methods, including RDS for hidden populations | Electronic health record linkage, web-based recruitment, social media, and mobile platforms | Active learning, including uncertainty sampling and query-by-committee; predictive recruitment models; adaptive surveillance | Probabilistic and AI-based methods, including active learning with mechanistic models [25] |
| Inference approach | Design-based inference using known selection probabilities | Often model-based, using post hoc adjustments | Algorithmic and predictive, focused on maximizing information gain | Combined design- and model-based inference that balances rigor and adaptability |
| Strengths | Strong theoretical guarantees of representativeness; valid causal inference | Expanded scale, efficiency, and real-time access | Prioritization of high-information cases; resource efficiency; dynamic adaptation | Enhanced forecasting, efficiency, and inferential validity; ability to bridge static and dynamic needs |
| Limitations | Static designs; inefficiency for large or heterogeneous data; high costs; declining response rates; difficulty studying hidden or hard-to-reach groups | Selection and coverage bias; reliance on post hoc modeling; digital divides, especially in LMICs | Risk of algorithmic bias amplification; black-box concerns; need for careful bias mitigation | Requires careful integration to avoid compounding bias |
| Bias handling | A priori, at the design stage through randomization and stratification | A posteriori, through modeling adjustments and weighting | Across the model lifecycle, including fairness-aware training | Systematic bias management across design, modeling, and algorithmic stages |
| Relevance to LMICs | Feasible when registries exist but constrained by incomplete frames and limited resources | Scalable but may exacerbate digital exclusion | High potential for efficiency in resource-constrained settings, including targeted case finding | Most promising when designed to address resource constraints while preserving equity, including pilots at facilities such as Dawanau |
| Examples from literature | National surveys and cohort studies [1,3] | Digital cohorts and web surveys [7] | Active learning in surveillance and targeted recruitment [11,12] | AI-mechanistic hybrid models for forecasting [25] |
| Suitability | Best suited for valid, unbiased population inference when sampling frames exist | Useful for rapid surveillance and hypothesis generation | Useful for real-time, resource-limited scenarios | Optimal pathway for combining efficiency, rigor, and equity |
| Aspects | Phase I: Classical probability-based sampling | Phase II: Digitally mediated sampling | Phase III: AI-enhanced adaptive sampling | Hybrid frameworks (recommended future) |
|---|---|---|---|---|
| Core emphasis | Design-based inference and randomization | Scalability and large-scale data access | Dynamic, information-driven selection | Integration of probabilistic foundations with adaptive algorithms |
| Key methods | Simple random, stratified, cluster, and multistage sampling; non-probability methods, including RDS for hidden populations | Electronic health record linkage, web-based recruitment, social media, and mobile platforms | Active learning, including uncertainty sampling and query-by-committee; predictive recruitment models; adaptive surveillance | Probabilistic and AI-based methods, including active learning with mechanistic models [25] |
| Inference approach | Design-based inference using known selection probabilities | Often model-based, using post hoc adjustments | Algorithmic and predictive, focused on maximizing information gain | Combined design- and model-based inference that balances rigor and adaptability |
| Strengths | Strong theoretical guarantees of representativeness; valid causal inference | Expanded scale, efficiency, and real-time access | Prioritization of high-information cases; resource efficiency; dynamic adaptation | Enhanced forecasting, efficiency, and inferential validity; ability to bridge static and dynamic needs |
| Limitations | Static designs; inefficiency for large or heterogeneous data; high costs; declining response rates; difficulty studying hidden or hard-to-reach groups | Selection and coverage bias; reliance on post hoc modeling; digital divides, especially in LMICs | Risk of algorithmic bias amplification; black-box concerns; need for careful bias mitigation | Requires careful integration to avoid compounding bias |
| Bias handling | A priori, at the design stage through randomization and stratification | A posteriori, through modeling adjustments and weighting | Across the model lifecycle, including fairness-aware training | Systematic bias management across design, modeling, and algorithmic stages |
| Relevance to LMICs | Feasible when registries exist but constrained by incomplete frames and limited resources | Scalable but may exacerbate digital exclusion | High potential for efficiency in resource-constrained settings, including targeted case finding | Most promising when designed to address resource constraints while preserving equity, including pilots at facilities such as Dawanau |
| Examples from literature | National surveys and cohort studies [1,3] | Digital cohorts and web surveys [7] | Active learning in surveillance and targeted recruitment [11,12] | AI-mechanistic hybrid models for forecasting [25] |
| Suitability | Best suited for valid, unbiased population inference when sampling frames exist | Useful for rapid surveillance and hypothesis generation | Useful for real-time, resource-limited scenarios | Optimal pathway for combining efficiency, rigor, and equity |
AI, artificial intelligence; RDS, respondent-driven sampling; LMICs, low-income and middle-income countries.