
FHIR for population health analytics uses Bulk Data IG $export to feed downstream analytics. Understanding the pipeline pattern prevents inefficient implementations.
Population health use cases
1. Risk stratification (readmission, sepsis, deterioration). 2. Care gap identification (missed screenings, medications). 3. Quality measure reporting. 4. Cohort analysis. 5. Outcomes tracking.
Bulk export patterns
1. Group-based export. Group/{id}/$export for population subsets. 2. Time-based incremental. $export?_since=<timestamp> for changes. 3. Resource-type filtered. $export?_type=Patient,Observation,Condition. 4. Full population export. For initial load or re-baseline.
Warehouse ingestion
1. NDJSON files → Spark, dbt, or cloud loaders. 2. Terminology snapshot joined at ingest. 3. Denormalized fact tables for BI. 4. Feature tables for ML.
Storage sizing (12M patients, mixed resources)
| Layer | Storage |
|---|---|
| Raw NDJSON (gzipped) | 40-60 GB |
| Warehouse tables | 100-200 GB |
| Feature stores | 10-50 GB |
Common pipeline failures
1. NDJSON chunk too large for loader. 2. Missing deleted[] handling. 3. No _since for incremental. 4. Runtime terminology $expand. 5. No time partitioning.
Vendor tooling
| Component | Options |
|---|---|
| FHIR export | HAPI, Aidbox, Medplum |
| Ingest | Spark, dbt, Databricks |
| Warehouse | BigQuery, Snowflake, Redshift |
| Feature store | Feast, dbt marts |
Data quality prerequisites
1. $validate pass rate >97%. 2. Reference integrity >99%. 3. Terminology compliance >98%. 4. Duplicate rate <2%.
Investment
1. Bulk export infrastructure. 2. Warehouse licensing. 3. Data engineering team. 4. BI tool licensing.
FHIR for population health is a solved discipline in 2026. The pipeline architecture is repeatable; execution quality determines value delivered.
