
Tuberculosis treatment records arrive in a separate clinical database and often lack the unique participant IDs used across six research cohorts, breaking the link between a patient's treatment and their original enrolment.
Without a shared identifier, manual reconciliation of treatment records against six studies was slow, error-prone and unscalable, leaving treatment outcomes disconnected from enrolment and screening data.
I built a two-stage probabilistic matcher — exact phone-hash lookup, then a demographic fuzzy match (rapidfuzz on names, filtered by age and gender) with weighted scoring — that reconciles treatment records to the right participant and syncs the result back to the central outcome database, with national-programme target tracking on top.
Treatment records are now linked automatically to their enrolment records and written back to the central database, and three Power BI dashboards track treatment progress against national-programme targets by site and region.