Skip Navigation
Skip to contents

Epidemiol Health : Epidemiology and Health

OPEN ACCESS
SEARCH
Search

Search

Page Path
HOME > Search
5 "Algorithm"
Filter
Filter
Article category
Publication year
Review
The evolution of sampling in epidemiology: from classical probability models to AI-enhanced recruitment and active learning
Sulaiman Abubakar Musa
Epidemiol Health. 2026;48:e2026019.   Published online May 5, 2026
DOI: https://doi.org/10.4178/epih.e2026019
  • 1,090 View
  • 28 Download
AbstractAbstract AbstractSummary PDF
Abstract
Sampling is a foundational element of epidemiological research because it determines the validity, efficiency, and generalizability of population-based inferences. Classical probability and non-probability sampling methods have long underpinned public health study design by providing robust frameworks for unbiased estimation and causal inference. However, the rapid expansion of digital health data, electronic health records, and large-scale online cohorts has exposed important limitations in these largely static sampling frameworks. This narrative review traces the evolution of sampling in epidemiology from traditional probability-based designs to contemporary artificial intelligence-enhanced recruitment strategies, with particular emphasis on active learning and adaptive sampling. By synthesizing classical statistical theory with emerging machine-learning approaches and recent methodological advances, this review argues that the future of epidemiological sampling lies in hybrid frameworks that preserve statistical rigor while leveraging algorithm-driven adaptability. Methodological opportunities, ethical risks, and implications for low-income and middle-income countries are critically examined.
Summary
Key Message
- Epidemiological sampling is evolving from static, design-based probability frameworks toward dynamic, AI-enhanced systems that improve efficiency and responsiveness in large-scale studies. - While classical sampling remains essential for inferential validity, integrating active learning and algorithmic recruitment enables researchers to prioritize high-information cases and navigate resource-constrained environments more effectively. - The future of the field lies in hybrid frameworks that combine statistical rigor with adaptive machine-learning strategies, provided these are supported by robust ethical governance to mitigate algorithmic bias and address health inequities.
Original Articles
Evaluation of gestational age by pregnancy outcomes and distribution of pregnancy-related codes in Korean claims data
Woo-Jung Kim, Yunha Noh, Yongtai Cho, Eun-Young Choi, HyunJoo Lim, Hyesung Lee, Ju-Young Shin
Epidemiol Health. 2026;48:e2026007.   Published online February 4, 2026
DOI: https://doi.org/10.4178/epih.e2026007
  • 4,751 View
  • 109 Download
AbstractAbstract AbstractSummary PDFSupplementary Material
Abstract
OBJECTIVES
This study aimed to evaluate a fixed-duration algorithm for gestational age (GA) estimation according to pregnancy outcomes and to describe the GA distribution of pregnancy-related codes in Korea.
METHODS
We included 351,055 pregnancy episodes (2019–2022) from linked data between the National Health Insurance Service and the Korea Immunization Registry Information System (KIRIS). GA from claims data was estimated by subtracting fixed durations from the delivery date (algorithm-based GA), and GA derived from KIRIS was defined as the gold standard. Accuracy was evaluated as the proportion of episodes in which the difference between the estimated GA and the reference standard fell within ±2 weeks. We described the distributions of the GA at which each prenatal test, pregnancy complication, and diagnostic code was recorded.
RESULTS
Algorithm-based GA estimation showed high accuracy for live births (92.2% within ±2 weeks) but markedly lower accuracy for non-live birth outcomes, including stillbirth (3.3%), termination (7.2%), spontaneous abortion (45.2%), and ectopic pregnancy (20.0%). In additional analyses aimed at identifying potential indicators for improving GA estimation, most events occurred within clinically expected timeframes, although some individual codes exhibited poor temporal alignment.
CONCLUSIONS
Algorithm-based GA estimation using claims data performed well for live births but demonstrated limited accuracy for non-live birth outcomes. Incorporating information from prenatal tests and pregnancy complications may enhance GA estimation.
Summary
Korean summary
본 연구는 국내 전국민 단위 청구자료와 예방접종 등록자료를 연계하여 고정 임신기간 기반 알고리즘의 재태연령 추정 정확도를 임신 결과별로 체계적으로 검증하였다. 출생아에서는 ±2주 기준 92.2%의 높은 정확도를 보였으나, 사산·인공임신중절·자연유산·자궁외임신 등 비출생 결과에서는 현저히 낮은 정확도를 보여 결과 유형에 따른 성능 이질성이 확인되었다. 산전검사 및 임신 합병증과 같은 시간 민감적 임상지표를 통합한 계층적 접근이 청구자료 기반 재태연령 추정의 타당도 향상에 기여할 수 있다.
Key Message
Using nationwide linked claims and immunization registry data in Korea, this study systematically validated the performance of a fixed-duration algorithm for gestational age estimation across pregnancy outcomes. While high accuracy was observed for live births (92.2% within ±2 weeks), substantially poorer performance was identified for non–live-birth outcomes, indicating marked outcome-specific heterogeneity. Integration of time-sensitive clinical indicators, including prenatal tests and pregnancy complications, may enhance the validity of gestational age estimation in administrative data research.
Predicting over-the-counter antibiotic use in rural Pune, India, using machine learning methods
Pravin Arun Sawant, Sakshi Shantanu Hiralkar, Yogita Purushottam Hulsurkar, Mugdha Sharad Phutane, Uma Satish Mahajan, Abhay Machindra Kudale
Epidemiol Health. 2024;46:e2024044.   Published online April 13, 2024
DOI: https://doi.org/10.4178/epih.e2024044
  • 15,374 View
  • 145 Download
AbstractAbstract PDFSupplementary Material
Abstract
OBJECTIVES
Over-the-counter (OTC) antibiotic use can cause antibiotic resistance, threatening global public health gains. To counter OTC use, this study used machine learning (ML) methods to identify predictors of OTC antibiotic use in rural Pune, India.
METHODS
The features of OTC antibiotic use were selected using stepwise logistic, lasso, random forest, XGBoost, and Boruta algorithms. Regression and tree-based models with all confirmed and tentatively important features were built to predict the use of OTC antibiotics. Five-fold cross-validation was used to tune the models’ hyperparameters. The final model was selected based on the highest area under the curve (AUROC) with a 95% confidence interval (CI) and the lowest log-loss.
RESULTS
In rural Pune, the prevalence of OTC antibiotic use was 35.9% (95% CI, 31.6 to 40.5). The perception that buying medicines directly from a medicine shop/pharmacy is useful, using antibiotics for eye-related complaints, more household members consuming antibiotics, and longer duration and higher doses of antibiotic consumption in rural blocks and other social groups were confirmed as important features by the Boruta algorithm. The final model was the XGBoost+Boruta model with 7 predictors (AUROC, 0.934; 95% CI, 0.891 to 0.978; log-loss, 0.279) log-loss.
CONCLUSIONS
XGBoost+Boruta, with 7 predictors, was the most accurate model for predicting OTC antibiotic use in rural Pune. Using OTC antibiotics for eye-related complaints, higher consumption of antibiotics and the perception that buying antibiotics directly from a medicine shop/pharmacy is useful were identified as key factors for planning interventions to improve awareness about proper antibiotic use.
Summary
Special Article
Identification of acute myocardial infarction and stroke events using the National Health Insurance Service database in Korea
Minsung Cho, Hyeok-Hee Lee, Jang-Hyun Baek, Kyu Sun Yum, Min Kim, Jang-Whan Bae, Seung-Jun Lee, Byeong-Keuk Kim, Young Ah Kim, JiHyun Yang, Dong Wook Kim, Young Dae Kim, Haeyong Pak, Kyung Won Kim, Sohee Park, Seng Chan You, Hokyou Lee, Hyeon Chang Kim
Epidemiol Health. 2024;46:e2024001.   Published online December 26, 2023
DOI: https://doi.org/10.4178/epih.e2024001
  • 25,863 View
  • 292 Download
  • 12 Web of Science
  • 10 Crossref
AbstractAbstract AbstractSummary PDF
Abstract
OBJECTIVES
The escalating burden of cardiovascular disease (CVD) is a critical public health issue worldwide. CVD, especially acute myocardial infarction (AMI) and stroke, is the leading contributor to morbidity and mortality in Korea. We aimed to develop algorithms for identifying AMI and stroke events from the National Health Insurance Service (NHIS) database and validate these algorithms through medical record review.
METHODS
We first established a concept and definition of “hospitalization episode,” taking into account the unique features of health claims-based NHIS database. We then developed first and recurrent event identification algorithms, separately for AMI and stroke, to determine whether each hospitalization episode represents a true incident case of AMI or stroke. Finally, we assessed our algorithms’ accuracy by calculating their positive predictive values (PPVs) based on medical records of algorithm-identified events.
RESULTS
We developed identification algorithms for both AMI and stroke. To validate them, we conducted retrospective review of medical records for 3,140 algorithm-identified events (1,399 AMI and 1,741 stroke events) across 24 hospitals throughout Korea. The overall PPVs for the first and recurrent AMI events were around 92% and 78%, respectively, while those for the first and recurrent stroke events were around 88% and 81%, respectively.
CONCLUSIONS
We successfully developed algorithms for identifying AMI and stroke events. The algorithms demonstrated high accuracy, with PPVs of approximately 90% for first events and 80% for recurrent events. These findings indicate that our algorithms hold promise as an instrumental tool for the consistent and reliable production of national CVD statistics in Korea.
Summary
Key Message
In this study, we developed algorithms to identify acute myocardial infarction (AMI) and stroke events from the Korean National Health insurance Service database. To validate them, we conducted retrospective review of medical records across 24 hospitals throughout Korea. The overall positive predictive values for the first and recurrent AMI events were around 92% and 78%, respectively, while those for the first and recurrent stroke events were around 88% and 81%, respectively.

Citations

Citations to this article as recorded by  
  • Cardiovascular Risk Among Stroke Survivors With Combustible and Electronic Cigarettes: A Nationwide Study in Korean Men
    Joonsang Yoo, Jimin Jeon, Minyoul Baik, Yun Young Choi, Jinkwon Kim
    Journal of the American Heart Association.2026;[Epub]     CrossRef
  • Blood pressure status and risk of cardiovascular disease in older adults aged 75+ without prior cardiovascular events: a nationwide cohort study
    Sangwon Choi, Kyung-Do Han, Kyung-Ho Yu, Byung-Chul Lee, Mi Sun Oh, Dae Young Cheon, Minwoo Lee
    European Journal of Preventive Cardiology.2026;[Epub]     CrossRef
  • Geographic disparities in EVT access and stroke mortality under universal health coverage in South Korea
    Jeehye Lee
    BMC Health Services Research.2026;[Epub]     CrossRef
  • Postacute sequelae of SARS-CoV-2 infection on ophthalmic diseases: a binational cohort study
    Jee Myung Yang, Hayeon Lee, Jaeyu Park, Min Kim, Ho-Seok Sa, Joo Yong Lee, Kyung Rim Sung, Francesco Branda, Krishna Prasad Acharya, Dong Keon Yon
    British Journal of Ophthalmology.2026; : bjo-2025-328412.     CrossRef
  • Constipation and risk of death and cardiovascular events in patients on hemodialysis
    Sang Cheol Park, Juyoung Jung, Young Eun Kwon, Song In Baeg, Dong-Jin Oh, Do Hyoung Kim, Young-Ki Lee, Hye Min Choi
    Kidney Research and Clinical Practice.2025; 44(1): 155.     CrossRef
  • Body Mass Index Changes and Femur Fracture Risk in Parkinson's Disease: National Cohort Study
    Sung‐Ho Ahn, Hye Sun Lee, Jun‐Hyuk Lee
    Journal of Cachexia, Sarcopenia and Muscle.2025;[Epub]     CrossRef
  • Confidence-linked and uncertainty-based staged framework for phenotype validation using large language models
    Sumin Lee, Hyeok-Hee Lee, Hokyou Lee, Kyu Sun Yum, Jang-Hyun Baek, Jaewon Khil, Jaeyong Lee, Sojung Shin, Minsung Cho, Na Yeon Ahn, Seng Chan You, Hyeon Chang Kim
    Journal of the American Medical Informatics Association.2025; 32(8): 1320.     CrossRef
  • Cardiovascular risk across blood pressure categories defined by the 2024 ESC and 2023 ESH hypertension guidelines: insights from a Korean nationwide cohort study
    Dae young Cheon, Kyung-do Han, Yeon Jung Lee, Jeen Hwa Lee, Myung Soo Park, Sook Jin Lee, Seongwoo Han, Jae Hyuk Choi, Minwoo Lee
    European Journal of Preventive Cardiology.2025;[Epub]     CrossRef
  • Cumulative Cardiovascular Health Score Through Young Adulthood and Cardiovascular and Kidney Outcomes in Midlife
    Jong Hyun Jhee, Kyoung Hwa Ha, Dasom Son, Hyeok-Hee Lee, Eun-Jin Kim, Hyeon Chang Kim, Hokyou Lee
    JAMA Cardiology.2025; 10(11): 1207.     CrossRef
  • Incidence and case fatality of stroke in Korea, 2011-2020
    Jenny Moon, Yeeun Seo, Hyeok-Hee Lee, Hokyou Lee, Fumie Kaneko, Sojung Shin, Eunji Kim, Kyu Sun Yum, Young Dae Kim, Jang-Hyun Baek, Hyeon Chang Kim
    Epidemiology and Health.2023; 46: e2024003.     CrossRef
Original Article
Identifying pregnancy episodes and estimating the last menstrual period using an administrative database in Korea: an application to patients with systemic lupus erythematosus
Yu-Seon Jung, Yeo-Jin Song, Jihyun Keum, Ju Won Lee, Eun Jin Jang, Soo-Kyung Cho, Yoon-Kyoung Sung, Sun-Young Jung
Epidemiol Health. 2024;46:e2024012.   Published online December 19, 2023
DOI: https://doi.org/10.4178/epih.e2024012
  • 20,383 View
  • 237 Download
  • 8 Web of Science
  • 8 Crossref
AbstractAbstract AbstractSummary PDFSupplementary Material
Abstract
OBJECTIVES
This study developed an algorithm for identifying pregnancy episodes and estimating the last menstrual period (LMP) in an administrative claims database and applied it to investigate the use of pregnancy-incompatible immunosuppressants among pregnant women with systemic lupus erythematosus (SLE).
METHODS
An algorithm was developed and applied to a nationwide claims database in Korea. Pregnancy episodes were identified using a hierarchy of pregnancy outcomes and clinically plausible periods for subsequent episodes. The LMP was estimated using preterm delivery, sonography, and abortion procedure codes. Otherwise, outcome-specific estimates were applied, assigning a fixed gestational age to the corresponding pregnancy outcome. The algorithm was used to examine the prevalence of pregnancies and utilization of pregnancy-incompatible immunosuppressants (cyclophosphamide [CYC]/mycophenolate mofetil [MMF]/methotrexate [MTX]) and non-steroidal anti-inflammatory drugs (NSAIDs) during pregnancy in SLE patients.
RESULTS
The pregnancy outcomes identified in SLE patients included live births (67%), stillbirths (2%), and abortions (31%). The LMP was mostly estimated with outcome-specific estimates for full-term births (92.3%) and using sonography procedure codes (54.7%) and preterm delivery diagnosis codes (37.9%) for preterm births. The use of CYC/MMF/MTX decreased from 7.6% during preconception to 0.2% at the end of pregnancy. CYC/MMF/MTX use was observed in 3.6% of women within 3 months preconception and 2.5% during 0-7 weeks of pregnancy.
CONCLUSIONS
This study presents the first pregnancy algorithm using a Korean administrative claims database. Although further validation is necessary, this study provides a foundation for evaluating the safety of medications during pregnancy using secondary databases in Korea, especially for rare diseases.
Summary
Korean summary
임산부의 약물 사용 안전성에 대한 근거 제공을 위해 실제 인구집단에서의 임신 중 약물 치료 안전성을 평가하는 청구자료 기반 연구가 중요하다. 본 연구에서는 국내 청구자료에 적용할 수 있는 임신 정의 및 임신 결과 조작적 정의 알고리즘을 개발하였다. 본 알고리즘은 임신 결과 간의 우선순위를 고려한 계층 구조를 활용하며, 조기 분만 및 초음파 검사 코드 등을 통해 최종 월경 기간을 추정하였다. 또한 알고리즘을 전신홍반루푸스 환자에 적용하여 유산, 사산 등의 유병률을 산출하고 임신 중 잠재적으로 부적절한 면역억제제 사용을 파악하여 국내 청구자료의 특성을 고려한 임신 중 약물 사용 연구의 기반을 마련하였다.
Key Message
Limited safety data for pregnant women prompted recent studies on medication during pregnancy using real-world databases. This study developed a tailored algorithm for Korean healthcare claims database, employing a hierarchy of pregnancy outcomes and incorporating pre-term delivery and sonography codes for last menstrual period estimation. Applied to systemic lupus erythematosus (SLE) patients, this study presented the prevalence and drug utilization pattern of pregnancy-incompatible immunosuppressants from preconception to pregnancy end, laying a foundation for further claims database studies on medication pregnancy safety.

Citations

Citations to this article as recorded by  
  • Immunosuppressant use and adverse pregnancy outcomes in women with systemic lupus erythematosus: a retrospective cohort study in Korea
    Yu-Seon Jung, Yeo-Jin Song, Eun Jin Jang, Soo-Kyung Cho, Yoon-Kyoung Sung, Sun-Young Jung
    Rheumatology.2026;[Epub]     CrossRef
  • Evaluation of gestational age by pregnancy outcomes and distribution of pregnancy-related codes in Korean claims data
    Woo-Jung Kim, Yunha Noh, Yongtai Cho, Eun-Young Choi, HyunJoo Lim, Hyesung Lee, Ju-Young Shin
    Epidemiology and Health.2026; 48: e2026007.     CrossRef
  • Evaluating and Refining Claims-Based Algorithms for Pregnancy Outcomes and Gestational Age Estimation in Korea
    Hee-Jin Kim, Haerin Cho, Sun-Young Jung, Ju-Young Shin, Seung-Ah Choe, Nam-Kyong Choi
    Journal of Korean Medical Science.2026;[Epub]     CrossRef
  • Development and Validation of Gestational Age Estimation Algorithms for Nonlive Births in Administrative Healthcare Databases
    Yongtai Cho, Eun-Young Choi, Hyesung Lee, Yunha Noh, Jung Yeol Han, Seung-Ah Choe, Hoon Kim, Ju-Young Shin
    Epidemiology.2026; 37(3): 336.     CrossRef
  • Identifying pregnancies in routinely collected health data: a scoping review of methods
    Paolo Mazzone, Rose Higgins, Siobhan O’Connor, Jenny Myers, Tjeerd van Staa, Victoria Palin
    BMC Medical Informatics and Decision Making.2026;[Epub]     CrossRef
  • Utilisation patterns of immunomodulators and pregnancy outcomes in systemic lupus erythematosus: Insights from Korean national data
    Yu-Seon Jung, Yeo-Jin Song, Hyeon Ji Lee, Eunji Kim, Soo-Kyung Cho, Yoon-Kyoung Sung, Sun-Young Jung
    Lupus.2025; 34(2): 140.     CrossRef
  • Narrative Review of National Health Insurance Claims Data for Gestational Age Calculation and Vaccine Safety Evaluation in Pregnancy
    Taemi Kim, Seung-Ah Choe
    Health Insurance Review & Assessment Service Research.2025; 5(2): 98.     CrossRef
  • Neurodevelopmental delays in children born after medically assisted reproduction: a national population cohort study
    Seung-Ah Choe, Eunseon Gwak, Juyoung Lee, Jung Hye Byeon, Ju-Young Shin, Seungbong Han, Jee Hyun Kim
    Journal of Neurodevelopmental Disorders.2025;[Epub]     CrossRef

Epidemiol Health : Epidemiology and Health
TOP