Skip to contents

Installation

# install.packages("remotes")
remotes::install_github("andrew-farrey/sysPrep")

Overview

sysPrep is an R package that provides formalized preprocessing methods for syndromic surveillance data from the National Syndromic Surveillance Program (NSSP) Electronic Surveillance System for the Early Notification of Community-based Epidemics (ESSENCE) API. Its methods focus on preparing ESSENCE data for real-world, applied public health surveillance.

sysPrep’s functions were developed through troubleshooting repeated false-positive overdose clusters back to duplicate ESSENCE visit records: a cause only visible in line-level data pulled via ESSENCE’s dataDetails API endpoint (Full Details). Its functions will not work on pre-aggregated ESSENCE data pulled using the tableBuilder or timeSeries API endpoints.

These steps originated in a statewide overdose surveillance system, but the value they add isn’t specific to overdose: sysPrep‘s methods broadly improve the quality of downstream ESSENCE data, particularly wherever ensuring accurate, distinct patient counts is an implicit part of the surveillance task. Results will vary by site: some sites’ Visit_ID may not support reliable deduplication, and for overdose surveillance specifically, where a larger share of EMS-responded overdoses involve refused transport, those patients are never captured in ED data at all regardless of preprocessing.

These methods are not required to perform case counting, cluster detection, or anomaly detection with ESSENCE data. Many surveillance questions tolerate the data quality issues sysPrep addresses without materially affecting conclusions. They are most valuable for small-count, high-impact case definitions, where external validity and minimizing false-positive clusters/anomalies matter most.

Raw ESSENCE data pulls can present four categories of data quality issues that bias case counts or distort cluster/anomaly detection if left unaddressed:

Problem Functions
Duplicate Records: multiple rows per visit due to visit date changes to the initial record, patient ID corrections, and patient class transitions dedupe(), summarize_duplicates(), classify_duplicates()
Non-Emergency Providers: facilities without EDs, and free-standing emergency departments (FSEDs) onboarded to ESSENCE with a non-emergency FacilityType, included in pulls filter_care_setting(), review_facility_ed_visits()
Mis-Submitted and Invisible Direct Admissions: some inpatient admissions are mistakenly submitted as unrelated to a preceding ED visit (a data quality artifact), and genuine direct admissions are excluded entirely from HasBeenE = 1 queries link_encounters()
Out-of-State and OTHER_REGION (unknown residence) Visits: these visits’ Region doesn’t match any in-state value, so ordinary region-scoped rollups (maps, county summary tables) silently exclude them with no explicit filter required, understating burden at the location where care was actually delivered assign_treating_geography(), assign_facility_geography()

sysPrep synthesizes these methods into a reproducible, documented pipeline.

Validated Data Sources

sysPrep’s functions have been validated against records pulled from the following NSSP ESSENCE data sources. Both sources return ED visit and inpatient admission records, so all exported functions apply to either:

ESSENCE datasource code Full name Used by
va_er, va_hosp Patient Location (Full Details), Facility Location (Full Details) dedupe(), summarize_duplicates(), classify_duplicates(), filter_care_setting(), review_facility_ed_visits(), link_encounters(), assign_treating_geography(), assign_facility_geography()

Column names and expected value formats (e.g., {SITE}_{REGION} region strings, HasBeenE/HasBeenAdmitted flags) reflect these two data sources. Pulls from other ESSENCE data sources may require column renaming before use.

Quick Start

library(sysPrep)

# Inspect raw data quality before preprocessing
summarize_duplicates(essence_raw)
classify_duplicates(essence_raw)

# Full preprocessing pipeline
clean <- essence_raw |>
  # Step 1: Remove duplicate records (retain most recently transmitted)
  dedupe(order_by = Arrived_Date_Time, keep = "last") |>
  # Step 2: Filter to ED and inpatient facilities; correct FSED types
  filter_care_setting(
    fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
  ) |>
  # Step 3: Assign treating facility geography to out-of-state visits
  # (writes region_hybrid/zip_code_hybrid; Region/ZipCode stay untouched)
  assign_treating_geography()

# Check for facility-level anomalies
clean |> review_facility_ed_visits(method = "both", date_col = Date)

For encounter linkage with a separate inpatient pull:

# Deduplicate both pulls
ed_clean        <- essence_ed        |> dedupe(order_by = Arrived_Date_Time)
inpatient_clean <- essence_inpatient |> dedupe(order_by = Arrived_Date_Time)

# Link into care episodes (captures direct admissions);
# one merged row per true encounter by default
episodes <- link_encounters(ed_clean, inpatient_clean)

# Distribution of care pathways
episodes |>
  dplyr::count(.patient_class_sequence, sort = TRUE)

For comparing residential vs. treating geography side by side, both geography functions write to new columns by default, so chaining them adds region_hybrid and region_facility without disturbing region itself (built fresh here rather than reusing clean above, since clean already ran assign_treating_geography() once as step 3 of its own pipeline):

geo_all <- essence_raw |>
  dedupe(order_by = Arrived_Date_Time, keep = "last") |>
  filter_care_setting(
    fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
  ) |>
  assign_treating_geography() |>
  assign_facility_geography()

geo_all |>
  dplyr::select(hospital_name, region, region_hybrid, region_facility)
# region: untouched original
# region_hybrid: residential geography, with treating geography substituted
#   only for out-of-state/unknown-residence visits
# region_facility: treating geography for every visit

Documentation

Full documentation including function reference pages and methodological vignettes is available at: https://andrew-farrey.github.io/sysPrep/

Vignettes: - Getting Started - Understanding and Resolving Duplicate Records - Linking ED and Inpatient Records - Geographic Attribution

  • Rnssp: NSSP ESSENCE API access, alerting algorithms, and syndromic surveillance utilities. sysPrep is designed to operate on data returned by Rnssp API calls; it offers an optional preprocessing layer between raw API output and case counting or cluster detection, for surveillance programs where that layer adds value.

Acknowledgements

sysPrep relies heavily on several packages whose authors deserve explicit credit:

  • janitor (Sam Firke): clean_names() is called on entry and exit of every function in the package. The column-name agnosticism that lets sysPrep accept both raw ESSENCE PascalCase and post-clean_names() snake_case is built entirely on this foundation.

  • dplyr, rlang, and tidyr (Hadley Wickham, Lionel Henry, and the tidyverse team): the data manipulation backbone and the inform() / warn() / abort() messaging infrastructure used throughout.

  • cli (Gábor Csárdi): the rich formatted output for the print.essence_dup_summary() and print.essence_dup_classified() S3 methods.

  • Rnssp (Gbedegnon Roseric Azondekon, Michael Sheppard, and the CDC BioSense team): the upstream package that handles NSSP ESSENCE API authentication and data retrieval. sysPrep would have no data to preprocess without it.

The methods formalized in sysPrep were developed through applied drug overdose surveillance and cluster detection work using NSSP ESSENCE data.

Citation

If you use sysPrep in published work, please cite the package directly:

citation("sysPrep")

Citation metadata is also available in CITATION.cff.

License

MIT © Andrew Farrey