Dev.to
5/11/2026

How a Data Pipeline and Machine Learning Can Modernize DNA Paternity Testing in Kenya
Short summary
Article proposes a machine learning data pipeline to modernize DNA paternity testing in Kenya, addressing slow turnaround times and lack of population-specific genetic databases. Current systems rely on manual STR analysis; researchers demonstrate supervised ML improves accuracy and speed. A staged pipeline is outlined—data ingestion, allele frequency reference construction, feature engineering, and model training using DNNs and gradient boosting.
- •ML can automate manual STR-based paternity analysis, improving speed and reducing human error in forensic analysis
- •Kenya lacks population-specific genetic reference data; most existing ML models trained on European and East Asian samples
- •Proposed 4-stage pipeline: standardized data ingestion, population-stratified allele frequency database, feature engineering (per-locus likelihood ratios), and model training with interpretable baselines (logistic regression, gradient boosting)
Generated with AI, which can make mistakes.
Is this a good recommendation for you?



