Skip to contents

Why Out-of-State Visits Exist

ESSENCE surveillance is facility-based: every patient treated at a participating facility is included in the data, regardless of where that patient lives. A facility near a state border or serving a metropolitan area that spans multiple states will routinely treat patients whose address of record is in a neighboring state. These patients generate valid, important clinical encounters: they may have overdosed at a Kentucky location even if their home is in Indiana or West Virginia.

Two additional sources of non-Kentucky geography appear in ESSENCE data:

OTHER_REGION records represent patients for whom a residential address is unavailable to the treating facility. This includes unhoused patients, patients who cannot communicate their address, and patients who decline to provide one. The OTHER_REGION value is not a geographic unit; it is a placeholder indicating that no residential geography can be assigned.

{SITE}_UNKNOWN (e.g. "KY_UNKNOWN") records represent visits at a known-site facility that did not transmit patient address fields. Unlike OTHER_REGION, this value carries the site prefix; it is ESSENCE’s placeholder for “the system knows this residence is missing,” as distinct from a residence that was received but does not map to any region. In practice, {SITE}_UNKNOWN is rare relative to OTHER_REGION: observed volumes in production Kentucky surveillance run in the single digits against tens of thousands of OTHER_REGION records over the same period, but because it carries the site prefix, it was silently misclassified as in-state and left unreassigned prior to this being caught. A bare NA in Region is treated as unknown residence too, though in practice ESSENCE typically fills in one of the placeholder values above rather than leaving the field empty.

The Cost of Discarding These Visits

The common response to out-of-state and unknown-residence records is to filter them out before analysis. This is understandable: residential geography is the expected denominator for incidence-based rates, but it introduces systematic bias. And it rarely requires an explicit filter() call to happen: any downstream step scoped to your jurisdiction’s regions, joining to a shapefile of in-state counties for a map, grouping by Region against a fixed list of in-state values for a summary table, silently drops rows whose Region doesn’t match, with no filter statement in the code to flag the decision for review.

Burden understatement. Facilities that treat a high volume of out-of-state or transient patients see their case counts reduced proportionally. A border facility that treats patients from three states appears to have a much lower burden than its actual clinical workload warrants.

Spatial bias. The exclusion rate is not uniform across facilities. Facilities in rural border communities and urban trauma centers serving mobile populations will have much higher exclusion rates than facilities in geographically isolated inland areas. When these exclusions are not accounted for, spatial comparisons across facilities are measuring a mix of true incidence and differential data completeness.

Secular trend distortion. If the composition of a facility’s patient population shifts over time, more transient patients, more border traffic, or a policy change affecting unhoused populations, the exclusion of OTHER_REGION records can create apparent trends in what is actually a stable underlying burden.

Reduced cluster/anomaly detection sensitivity. Hospital-level cluster detection sees every visit treated at a facility, regardless of patient residence. Region- or zip-level detection using unmodified Region doesn’t: out-of-state and unknown-residence visits either fall outside the surveillance area entirely or get binned to a region unrelated to where the activity occurred. A border facility’s own region can show a materially smaller signal than that facility’s true hospital-level volume, purely from this dropout, not from any real difference in incidence. Reassigning those visits to the treating facility’s region closes most of that gap, bringing region-level sensitivity close to hospital-level, though not identical when a region contains more than one hospital (see ?assign_facility_geography’s own discussion of hospital-level vs. region-level cluster detection for that separate tradeoff).

This is also why sysPrep treats the fix as attribution rather than exclusion-avoidance alone: EMS-based systems (e.g., ODMAP, Biospatial) report incidence at the location where care was rendered, not the patient’s jurisdiction of residence. Assigning treating facility geography brings ESSENCE data into that same incidence-based frame, which matters for applied, rapid-response surveillance where the operative question is where a case was treated, not which jurisdiction is responsible for the follow-up.

sysPrep provides two functions for retaining these visits by attributing them to the geography of the treating facility. The choice of which function to use depends on the research question.

What Region Actually Represents

Is Region an authoritative county boundary? No. Region is ESSENCE’s standardized construct for enabling consistent sub-state reporting across data sources: a maintained, many-to-one zip-code-to-region lookup table, not a live geocode computed from each record. By default, a zip code is assigned to a region using its geographic centroid, but individual assignments can be, and routinely are, overridden by site or state administrators to better reflect where the bulk of a zip code’s population actually lives. The same logic applies to HospitalRegion: assigned from the facility’s zip centroid by default, but overridable to the facility’s listed county. Because of this approximation, results reported by Region should not be construed as the authoritative count for that county; they are the best available standardized proxy, and the mapping in effect today is not necessarily the mapping that applied when a record was first received.

This matters for sysPrep because both geography functions operate entirely on Region/HospitalRegion as received; they reassign which existing value applies to a row, they do not independently geocode or validate the underlying zip-to-county mapping.

It also explains why two different functions exist at all, rather than one that always attributes to “the” geography. ESSENCE itself distinguishes between patient location and facility location data sources:

  • In a patient location pull (e.g., va_er), the display criterion is the patient’s residence: the data show who, from a given jurisdiction, presented for care anywhere.
  • In a facility location pull (e.g., va_hosp), the display criterion is the treating facility: the data show everyone who presented at an in-jurisdiction facility, combining residents and visitors alike.

essence_raw and essence_clean model a patient location pull, so Region reflects patient residence out of the box. assign_treating_geography() preserves that patient-location semantics for the in-state majority and only substitutes facility geography as a fallback. assign_facility_geography() instead converts the dataset to facility-location semantics uniformly, useful when your source pull, or your research question, is facility-location in nature to begin with.

The Region Field Format

ESSENCE encodes geography in the format {SITE}_{REGION}, where SITE is the NSSP Site Short Name (e.g., "KY") and REGION is the ESSENCE Region, a county name derived from a zip-code-to-county lookup table maintained by ESSENCE (e.g., "Jefferson"). Some site names contain additional underscores (e.g., multi-word state abbreviations used in certain NSSP configurations), so the site prefix is defined as all characters before the last underscore in the string.

Both geography functions use the same detection logic: - A visit is classified as in-state when Region begins with paste0(site, "_") (e.g., "KY_" for site = "KY") and is not the known-site unknown-residence placeholder below. - A visit is classified as out-of-state or unknown when Region begins with a different prefix, is exactly "OTHER_REGION", is exactly paste0(site, "_UNKNOWN") (e.g., "KY_UNKNOWN"), or is NA. The "{site}_UNKNOWN" form is checked explicitly rather than relying on the prefix mismatch, since it carries the site prefix and would otherwise be (incorrectly) classified as in-state.

# Illustrate the detection logic on a few example values
region_examples <- c("KY_Jefferson", "TN_Davidson", "OH_Hamilton",
                     "OTHER_REGION", "KY_UNKNOWN", NA_character_)

data.frame(
  region   = region_examples,
  in_state = startsWith(
    tidyr::replace_na(region_examples, ""),
    "KY_"
  ) &
    !is.na(region_examples) &
    region_examples != "KY_UNKNOWN"
)
#>         region in_state
#> 1 KY_Jefferson     TRUE
#> 2  TN_Davidson    FALSE
#> 3  OH_Hamilton    FALSE
#> 4 OTHER_REGION    FALSE
#> 5   KY_UNKNOWN    FALSE
#> 6         <NA>    FALSE

Two Functions, Two Philosophies

sysPrep provides two geographic attribution functions that differ fundamentally in which rows are modified:

assign_treating_geography() assign_facility_geography()
Default output column region_hybrid (new column; Region untouched) region_facility (new column; Region untouched)
Rows reassigned Out-of-state + OTHER_REGION + NA only All rows
In-state patient geography Reflected as-is in region_hybrid Overwritten with facility geography
Research question Where did this care take place? (for non-residents only) Where did all care take place?
Resulting column meaning Mixed: residential for in-state, treating for out-of-state Treating facility county for everyone

Both default to writing a new column rather than overwriting Region in place; Region itself is never touched unless you explicitly opt into that with overwrite = TRUE (covered at the end of this vignette). This means calling both functions back to back gives you all three geographies side by side with zero extra arguments:

geo_all <- essence_raw |>
  dedupe(order_by = Arrived_Date_Time, keep = "last") |>
  filter_care_setting(
    fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
  ) |>
  assign_treating_geography() |>
  assign_facility_geography()
#> The following `FacilityType` values are not in `keep_types` and will be excluded:
#>   - Primary Care
#>   - Medical Specialty
#> 27 of 160 visits (16.9%) identified as out-of-state or OTHER_REGION and assigned treating facility geography in `region_hybrid`/`zip_code_hybrid`.
#> Facility geography applied to all 160 visits in `region_facility`/`zip_code_facility`.

geo_all |>
  dplyr::select(hospital_name, region, region_hybrid, region_facility) |>
  head(10)
#> # A tibble: 10 × 4
#>    hospital_name          region       region_hybrid region_facility
#>    <chr>                  <chr>        <chr>         <chr>          
#>  1 Central Medical Center TN_Davidson  KY_Jefferson  KY_Jefferson   
#>  2 Central Medical Center KY_Kenton    KY_Kenton     KY_Jefferson   
#>  3 Central Medical Center KY_Campbell  KY_Campbell   KY_Jefferson   
#>  4 Central Medical Center OH_Hamilton  KY_Jefferson  KY_Jefferson   
#>  5 Central Medical Center OH_Franklin  KY_Jefferson  KY_Jefferson   
#>  6 Central Medical Center KY_McCracken KY_McCracken  KY_Jefferson   
#>  7 Central Medical Center KY_Daviess   KY_Daviess    KY_Jefferson   
#>  8 Central Medical Center KY_Daviess   KY_Daviess    KY_Jefferson   
#>  9 Central Medical Center NA           KY_Jefferson  KY_Jefferson   
#> 10 Central Medical Center KY_Campbell  KY_Campbell   KY_Jefferson

region is the untouched original for every row; region_hybrid matches it for in-state visits and substitutes treating geography for the rest; region_facility substitutes treating geography for every row, including the in-state majority. One pass over the data, three usable geographies, no preserve_original_geographies, no manual copying.

This distinction matters for how the resulting data should be interpreted and which analyses it supports.

assign_treating_geography(): Selective Hybrid Reassignment

assign_treating_geography() is a targeted intervention on problem rows. It identifies visits where the patient geography is unavailable or out-of-area and writes the treating facility’s county, for those visits only, to region_hybrid. region is left exactly as received from ESSENCE for every row, in-state and out-of-state alike.

region_hybrid is a hybrid column: it holds residential geography for patients who provided an in-state address, and treating facility geography for patients who could not or did not. This hybrid is appropriate for area-based incidence analyses where you want to count as many visits as possible within your surveillance area, and where the distinction between resident and non-resident burden is not the primary analytical question.

# Start from deduplicated, filtered data
ed_clean <- essence_raw |>
  dedupe(order_by = Arrived_Date_Time, keep = "last") |>
  filter_care_setting(
    fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
  )
#> The following `FacilityType` values are not in `keep_types` and will be excluded:
#>   - Primary Care
#>   - Medical Specialty

# Selective reassignment: only out-of-state and OTHER_REGION rows get a
# reassigned value in region_hybrid; region itself is never touched
treating_geo <- ed_clean |>
  assign_treating_geography(site = "KY")
#> 27 of 160 visits (16.9%) identified as out-of-state or OTHER_REGION and
#> assigned treating facility geography in `region_hybrid`/`zip_code_hybrid`.
# How many visits were reassigned?
dplyr::count(treating_geo, .out_of_state)
#> # A tibble: 2 × 2
#>   .out_of_state     n
#>   <lgl>         <int>
#> 1 FALSE           133
#> 2 TRUE             27
# Compare original region against the hybrid result for reassigned rows;
# no preserve_original_geographies needed, since region was never modified
treating_geo |>
  dplyr::filter(.out_of_state) |>
  dplyr::select(hospital_name, region, region_hybrid) |>
  head(10)
#> # A tibble: 10 × 3
#>    hospital_name          region       region_hybrid
#>    <chr>                  <chr>        <chr>        
#>  1 Central Medical Center TN_Davidson  KY_Jefferson 
#>  2 Central Medical Center OH_Hamilton  KY_Jefferson 
#>  3 Central Medical Center OH_Franklin  KY_Jefferson 
#>  4 Central Medical Center NA           KY_Jefferson 
#>  5 Central Medical Center OTHER_REGION KY_Jefferson 
#>  6 Central Medical Center TN_Shelby    KY_Jefferson 
#>  7 Central Medical Center OH_Hamilton  KY_Jefferson 
#>  8 North County Hospital  TN_Shelby    KY_Kenton    
#>  9 North County Hospital  WV_Cabell    KY_Kenton    
#> 10 North County Hospital  NA           KY_Kenton

The .out_of_state column (logical) marks every row where region_hybrid differs from region. Rows where .out_of_state = FALSE have region_hybrid == region: the patient’s residential county, unchanged.

If you’d rather overwrite region/zip_code in place instead of getting a region_hybrid/zip_code_hybrid column (reproducing the only behavior this function had before new_region_col/new_zip_col existed), pass their own names explicitly and opt in with overwrite = TRUE:

# Opt-in in-place overwrite, with an audit trail of what was replaced
treating_overwrite <- ed_clean |>
  assign_treating_geography(
    site                           = "KY",
    new_region_col                 = "region",
    new_zip_col                    = "zip_code",
    overwrite                      = TRUE,
    preserve_original_geographies  = TRUE
  )
#> 27 of 160 visits (16.9%) identified as out-of-state or OTHER_REGION and
#> assigned treating facility geography in `region`/`zip_code`.

treating_overwrite |>
  dplyr::filter(.out_of_state) |>
  dplyr::select(hospital_name, original_region, region) |>
  head(10)
#> # A tibble: 10 × 3
#>    hospital_name          original_region region      
#>    <chr>                  <chr>           <chr>       
#>  1 Central Medical Center TN_Davidson     KY_Jefferson
#>  2 Central Medical Center OH_Hamilton     KY_Jefferson
#>  3 Central Medical Center OH_Franklin     KY_Jefferson
#>  4 Central Medical Center NA              KY_Jefferson
#>  5 Central Medical Center OTHER_REGION    KY_Jefferson
#>  6 Central Medical Center TN_Shelby       KY_Jefferson
#>  7 Central Medical Center OH_Hamilton     KY_Jefferson
#>  8 North County Hospital  TN_Shelby       KY_Kenton   
#>  9 North County Hospital  WV_Cabell       KY_Kenton   
#> 10 North County Hospital  NA              KY_Kenton

preserve_original_geographies only has an effect in this overwrite = TRUE mode: when it applies, original_region/original_zip_code are added before overwriting, retaining pre-reassignment values as an audit trail. For in-state visits, original_region is NA because no change was made. In the default (new-column) mode above, this parameter is unnecessary and has no effect, since region was never modified in the first place.

When to use assign_treating_geography():

  • Incidence estimation where you want residential geography for most patients and facility geography only as a fallback for patients without valid residential addresses.
  • Trend analyses where you need to retain border-facility visits without attributing all visits to facility location.
  • Sensitivity analyses: compare results with and without OTHER_REGION visits using the .out_of_state flag to filter.

assign_facility_geography(): Universal Full Reassignment

assign_facility_geography() takes a fundamentally different approach: every row gets a facility-geography value in region_facility, regardless of whether the patient was in-state, out-of-state, or unknown. region itself is untouched, exactly as with assign_treating_geography(). region_facility never reflects where patients live; it always reflects where they received care.

# Universal reassignment: ALL rows get a facility geography value
facility_geo <- ed_clean |> assign_facility_geography()
#> Facility geography applied to all 160 visits in
#> `region_facility`/`zip_code_facility`.
# region_facility vs. the untouched original, for every row
facility_geo |>
  dplyr::select(hospital_name, hospital_region, region, region_facility) |>
  head(10)
#> # A tibble: 10 × 4
#>    hospital_name          hospital_region region       region_facility
#>    <chr>                  <chr>           <chr>        <chr>          
#>  1 Central Medical Center KY_Jefferson    TN_Davidson  KY_Jefferson   
#>  2 Central Medical Center KY_Jefferson    KY_Kenton    KY_Jefferson   
#>  3 Central Medical Center KY_Jefferson    KY_Campbell  KY_Jefferson   
#>  4 Central Medical Center KY_Jefferson    OH_Hamilton  KY_Jefferson   
#>  5 Central Medical Center KY_Jefferson    OH_Franklin  KY_Jefferson   
#>  6 Central Medical Center KY_Jefferson    KY_McCracken KY_Jefferson   
#>  7 Central Medical Center KY_Jefferson    KY_Daviess   KY_Jefferson   
#>  8 Central Medical Center KY_Jefferson    KY_Daviess   KY_Jefferson   
#>  9 Central Medical Center KY_Jefferson    NA           KY_Jefferson   
#> 10 Central Medical Center KY_Jefferson    KY_Campbell  KY_Jefferson
# Confirm: region_facility == hospital_region for every row
all(facility_geo$region_facility == facility_geo$hospital_region)
#> [1] TRUE

Unlike assign_treating_geography(), where only out-of-state/unknown rows differ between region and region_hybrid, here every row differs between region and region_facility; that’s the “universal” part.

Is this only renaming HospitalRegion to Region? Mechanically, close: assign_facility_geography() writes HospitalRegion and/or HospitalZip to region_facility/zip_code_facility for all rows, depending on which columns are present in the pull. Not all analysts include both; zip-based geography carries its own methodological considerations around spatial resolution and boundary alignment that may not be appropriate for every analysis. When only HospitalRegion is present, only region is reassigned. Some practitioners have applied this rename informally as an ad hoc pipeline step. sysPrep formalizes that practice: documents the intent, handles whichever geography types are available, and writes to a new column by default so the original values are always available for audit or comparison without any extra argument. assign_treating_geography(), the hybrid selective approach, has no simple column-rename equivalent, which is what motivated formalizing both methods in the same package.

Hospital-level spatial clustering (point-based) concentrates observations at individual facility locations and can identify which specific facilities are experiencing elevated burden, including spikes that would be diluted or invisible at the regional level. However, hospital-level scan statistics generally do not explicitly model residual spatial autocorrelation or facility-specific reporting patterns as nuisance structure. In prospective surveillance, scanning windows may therefore interact with hospital density, catchment patterns, and data submission variability in ways that can produce apparent clusters independent of true changes in underlying incidence. This concern is more pronounced in urban areas, where hospitals are denser and submission patterns may vary across nearby facilities. Reassigning all visits to HospitalRegion before aggregation reduces facility-level noise by aggregating across treating facilities within a region, allowing spacetime permutation methods and other scan statistics to treat a rural county with one hospital comparably to an urban county with many. Region-level binning does not eliminate spatial autocorrelation concerns, but it reduces the influence of facility-specific variability that is more common in high-density areas.

A second advantage is logistical: hospital-level spatial clustering requires latitude/longitude coordinates from the NSSP Master Facility Table, an additional lookup outside the standard ESSENCE data pull. Region-level attribution achieves a close approximation of hospital-level cluster detection using HospitalRegion, a field returned without any coordinate lookup, provided it’s included in the pull: HospitalRegion isn’t guaranteed on every ESSENCE pull by default, only when it’s part of your site’s default returned columns or explicitly requested via the API’s &field= parameter. The same logistical advantage also supports a distinct, simpler use case: total treated volume by urban vs. rural county, with no interest in which specific hospital saw each visit at all, not an approximation of hospital-level detection, just county-level counting using a field you already requested rather than a coordinate lookup you’d have to add.

What does a county with no ESSENCE-participating hospital show? 0, always: not a small or uncertain number, a hard zero, regardless of how much true burden that county’s residents actually generate. This is a direct consequence of ESSENCE being facility-based: region_facility can only ever reflect counties containing a facility that submits to ESSENCE. A county with real burden but no local ED will never contribute a count under this method, no matter how the underlying incidence changes. This is worth stating plainly rather than discovering it from a suspiciously flat map.

The two approaches are complementary: hospital-based clustering offers facility-level resolution; region-based clustering offers geographic equity across areas of differing hospital density. The choice depends on whether the surveillance question is “which facility?” or “which area?”.

When to use assign_facility_geography():

  • Analyses framed around treating facility location rather than patient residence: “How many cases were treated in Jefferson County?” rather than “How many residents of Jefferson County were treated?”
  • Facility service area analyses where you want all visits attributed to the county where the facility is located.
  • Spatial cluster detection scoped to treating location (e.g., identifying counties with high overdose treatment burden regardless of patient origin).
  • Cross-state facility comparisons where residential geography is an inappropriate common denominator.
  • Comparing total treated volume across urban and rural counties without differentiating which hospital saw each visit, useful when a facility-to-county lookup table isn’t available but HospitalRegion/ HospitalZip already are.

The Core Distinction: What Do the Output Columns Mean?

The choice between the two functions is a choice about what geography your analysis should be scoped to:

region_hybrid (from assign_treating_geography()): residential geography for in-state patients + treating facility geography for out-of-state and unknown-residence patients.

This is a hybrid with mixed semantics. It is most useful when you want to maximize the number of visits attributed to your surveillance area while preserving the residential geography of the majority.

region_facility (from assign_facility_geography()): treating facility geography for all patients.

This is semantically uniform: region_facility means the same thing for every row. It is most useful when the analytical unit is where care was delivered, not where patients live.

region itself, meanwhile, is never touched by either function unless you explicitly opt into overwrite = TRUE, so it’s always available as the untouched original, regardless of which (or both) of these you compute. Neither function is universally correct. The choice should be documented in your analysis methods and driven by the research question.

Decision Guide

Research question Function Rationale
County-level incidence rate (retain most OOS visits) assign_treating_geography() Preserves residential geography for in-state majority; uses facility geography as fallback
Which counties have highest treatment burden? assign_facility_geography() All visits attributed to treating county uniformly
Where do patients live who are treated at this facility? Neither; use Region directly Residential geography is the question; don’t reassign
Spatial cluster detection by facility county assign_facility_geography() Uniform attribution aligns with facility-based cluster geometry
Sensitivity analysis: effect of OOS exclusion assign_treating_geography() Use .out_of_state flag to toggle OOS visits in/out
Trend analysis, mixed state border facility assign_treating_geography() Avoids flattening in-state residential variation

Output Columns: New by Default, Overwrite as an Opt-In

Every example above wrote to a new column (region_hybrid/region_facility) and left region/zip_code alone. That’s the default for a reason: it’s non-destructive, so calling either function is always safe to try, safe to re-run, and safe to chain: there’s no risk of losing the original values just by calling the function.

new_region_col/new_zip_col control the target column and default to "region_hybrid"/"zip_code_hybrid" (assign_treating_geography()) or "region_facility"/"zip_code_facility" (assign_facility_geography()). Pass NULL to either to skip that geography type entirely, e.g. new_zip_col = NULL to reassign region only.

If the target you name already exists as a column, whether that’s region_col/zip_col itself, or any other existing column, the function aborts rather than silently overwriting it, unless you set overwrite = TRUE. This is what makes reproducing the original in-place-overwrite behavior an explicit, visible choice rather than the default:

essence_clean |>
  assign_treating_geography(
    new_region_col = "region",
    new_zip_col    = "zip_code",
    overwrite      = TRUE
  )

preserve_original_geographies = TRUE only has an effect in this overwrite = TRUE mode: it adds original_region/original_zip_code columns holding the pre-overwrite values before they’re replaced. In the default new-column mode, it has no effect and is unnecessary, since region was never modified in the first place; it already is the original.

For assign_treating_geography() in overwrite mode, original_region/ original_zip_code contain pre-reassignment values only for rows where geography changed (.out_of_state = TRUE); for in-state visits they are NA. For assign_facility_geography() in overwrite mode, they contain pre-reassignment values for every row, because every row is reassigned.