
A near point-of-care TB testing study at community health centres collects QR-coded field data across screening, specimen and laboratory forms that must be reconciled into three clinical databases and monitored against a study protocol.
Field testing generates messy, duplicate-prone data dumps spread across many forms and sites, with no automatic way to catch quality problems or track progress against per-site testing targets.
I built an end-to-end pipeline that deduplicates and uploads field dumps into three REDCap projects (thirteen forms), then runs a configurable data-quality engine that scores records on multiple dimensions, persists query lifetimes, and builds region-aware daily and monthly testing targets using a working-day calendar with Cameroon public holidays.
Field data flows from collection to clean, quality-scored REDCap records automatically, with daily tracking reports and per-site targets that let the study monitor testing coverage and data quality in near-real time.