Clinical data engineering — near point-of-care TB testing pipeline

Near point-of-care TB testing: an automated dump-to-REDCap pipeline with built-in data-quality scoring

reliable, continuously quality-checked field data from community testing sites

A health worker performing a point-of-care check at a community health centre.
Near point-of-care TB testing at community health centres — field dumps to clean, quality-scored records (representative photo).

A near point-of-care TB testing study at community health centres collects QR-coded field data across screening, specimen and laboratory forms that must be reconciled into three clinical databases and monitored against a study protocol.

The problem

Field testing generates messy, duplicate-prone data dumps spread across many forms and sites, with no automatic way to catch quality problems or track progress against per-site testing targets.

What I did

I built an end-to-end pipeline that deduplicates and uploads field dumps into three REDCap projects (thirteen forms), then runs a configurable data-quality engine that scores records on multiple dimensions, persists query lifetimes, and builds region-aware daily and monthly testing targets using a working-day calendar with Cameroon public holidays.

3 REDCap projects
Orchestrated from one pipeline
13 forms
Screening, specimen & lab
8 test types
Across 3 diagnostic categories
4-level
Data-quality priority taxonomy
data-qualityREDCapdeduplicationschedulingSQLitePython
Result

Field data flows from collection to clean, quality-scored REDCap records automatically, with daily tracking reports and per-site targets that let the study monitor testing coverage and data quality in near-real time.

LET'S TALK

Ready to get started?

Tell me about your project, question or goal. I reply within 24 hours, in English or French.

Get in touch