
Geographic Attribution for Out-of-State and Unknown Residence Visits
Source:vignettes/geography-assignment.Rmd
geography-assignment.RmdWhy Out-of-State Visits Exist
ESSENCE surveillance is facility-based: every patient treated at a participating facility is included in the data, regardless of where that patient lives. A facility near a state border or serving a metropolitan area that spans multiple states will routinely treat patients whose address of record is in a neighboring state. These patients generate valid, important clinical encounters: they may have overdosed at a Kentucky location even if their home is in Indiana or West Virginia.
Two additional sources of non-Kentucky geography appear in ESSENCE data:
OTHER_REGION records represent patients
for whom a residential address is unavailable to the treating facility.
This includes unhoused patients, patients who cannot communicate their
address, and patients who decline to provide one. The
OTHER_REGION value is not a geographic unit; it is a
placeholder indicating that no residential geography can be
assigned.
{SITE}_UNKNOWN
(e.g. "KY_UNKNOWN") records represent visits at a
known-site facility that did not transmit patient address fields. Unlike
OTHER_REGION, this value carries the site prefix; it is
ESSENCE’s placeholder for “the system knows this residence is missing,”
as distinct from a residence that was received but does not map to any
region. In practice, {SITE}_UNKNOWN is rare relative to
OTHER_REGION: observed volumes in production Kentucky
surveillance run in the single digits against tens of thousands of
OTHER_REGION records over the same period, but because it
carries the site prefix, it was silently misclassified as in-state and
left unreassigned prior to this being caught. A bare NA in
Region is treated as unknown residence too, though in
practice ESSENCE typically fills in one of the placeholder values above
rather than leaving the field empty.
The Cost of Discarding These Visits
The common response to out-of-state and unknown-residence records is
to filter them out before analysis. This is understandable: residential
geography is the expected denominator for incidence-based rates, but it
introduces systematic bias. And it rarely requires an explicit
filter() call to happen: any downstream step scoped to your
jurisdiction’s regions, joining to a shapefile of in-state counties for
a map, grouping by Region against a fixed list of in-state
values for a summary table, silently drops rows whose
Region doesn’t match, with no filter statement in the code
to flag the decision for review.
Burden understatement. Facilities that treat a high volume of out-of-state or transient patients see their case counts reduced proportionally. A border facility that treats patients from three states appears to have a much lower burden than its actual clinical workload warrants.
Spatial bias. The exclusion rate is not uniform across facilities. Facilities in rural border communities and urban trauma centers serving mobile populations will have much higher exclusion rates than facilities in geographically isolated inland areas. When these exclusions are not accounted for, spatial comparisons across facilities are measuring a mix of true incidence and differential data completeness.
Secular trend distortion. If the composition of a
facility’s patient population shifts over time, more transient patients,
more border traffic, or a policy change affecting unhoused populations,
the exclusion of OTHER_REGION records can create apparent
trends in what is actually a stable underlying burden.
Reduced cluster/anomaly detection sensitivity.
Hospital-level cluster detection sees every visit treated at a facility,
regardless of patient residence. Region- or zip-level detection using
unmodified Region doesn’t: out-of-state and
unknown-residence visits either fall outside the surveillance area
entirely or get binned to a region unrelated to where the activity
occurred. A border facility’s own region can show a materially smaller
signal than that facility’s true hospital-level volume, purely from this
dropout, not from any real difference in incidence. Reassigning those
visits to the treating facility’s region closes most of that gap,
bringing region-level sensitivity close to hospital-level, though not
identical when a region contains more than one hospital (see
?assign_facility_geography’s own discussion of
hospital-level vs. region-level cluster detection for that separate
tradeoff).
This is also why sysPrep treats the fix as attribution
rather than exclusion-avoidance alone: EMS-based systems (e.g., ODMAP,
Biospatial) report incidence at the location where care was rendered,
not the patient’s jurisdiction of residence. Assigning treating facility
geography brings ESSENCE data into that same incidence-based frame,
which matters for applied, rapid-response surveillance where the
operative question is where a case was treated, not which jurisdiction
is responsible for the follow-up.
sysPrep provides two functions for retaining these
visits by attributing them to the geography of the treating facility.
The choice of which function to use depends on the research
question.
What Region Actually Represents
Is
Regionan authoritative county boundary? No.Regionis ESSENCE’s standardized construct for enabling consistent sub-state reporting across data sources: a maintained, many-to-one zip-code-to-region lookup table, not a live geocode computed from each record. By default, a zip code is assigned to a region using its geographic centroid, but individual assignments can be, and routinely are, overridden by site or state administrators to better reflect where the bulk of a zip code’s population actually lives. The same logic applies toHospitalRegion: assigned from the facility’s zip centroid by default, but overridable to the facility’s listed county. Because of this approximation, results reported byRegionshould not be construed as the authoritative count for that county; they are the best available standardized proxy, and the mapping in effect today is not necessarily the mapping that applied when a record was first received.
This matters for sysPrep because both geography
functions operate entirely on
Region/HospitalRegion as received; they
reassign which existing value applies to a row, they do not
independently geocode or validate the underlying zip-to-county
mapping.
It also explains why two different functions exist at all, rather than one that always attributes to “the” geography. ESSENCE itself distinguishes between patient location and facility location data sources:
- In a patient location pull (e.g.,
va_er), the display criterion is the patient’s residence: the data show who, from a given jurisdiction, presented for care anywhere. - In a facility location pull (e.g.,
va_hosp), the display criterion is the treating facility: the data show everyone who presented at an in-jurisdiction facility, combining residents and visitors alike.
essence_raw and essence_clean model a
patient location pull, so Region reflects patient residence
out of the box. assign_treating_geography() preserves that
patient-location semantics for the in-state majority and only
substitutes facility geography as a fallback.
assign_facility_geography() instead converts the dataset to
facility-location semantics uniformly, useful when your source pull, or
your research question, is facility-location in nature to begin
with.
The Region Field Format
ESSENCE encodes geography in the format {SITE}_{REGION},
where SITE is the NSSP Site Short Name (e.g.,
"KY") and REGION is the ESSENCE Region, a
county name derived from a zip-code-to-county lookup table maintained by
ESSENCE (e.g., "Jefferson"). Some site names contain
additional underscores (e.g., multi-word state abbreviations used in
certain NSSP configurations), so the site prefix is defined as all
characters before the last underscore in the
string.
Both geography functions use the same detection logic: - A visit is
classified as in-state when Region begins
with paste0(site, "_") (e.g., "KY_" for
site = "KY") and is not the known-site unknown-residence
placeholder below. - A visit is classified as out-of-state or
unknown when Region begins with a different
prefix, is exactly "OTHER_REGION", is exactly
paste0(site, "_UNKNOWN") (e.g., "KY_UNKNOWN"),
or is NA. The "{site}_UNKNOWN" form is checked
explicitly rather than relying on the prefix mismatch, since it carries
the site prefix and would otherwise be (incorrectly) classified as
in-state.
# Illustrate the detection logic on a few example values
region_examples <- c("KY_Jefferson", "TN_Davidson", "OH_Hamilton",
"OTHER_REGION", "KY_UNKNOWN", NA_character_)
data.frame(
region = region_examples,
in_state = startsWith(
tidyr::replace_na(region_examples, ""),
"KY_"
) &
!is.na(region_examples) &
region_examples != "KY_UNKNOWN"
)
#> region in_state
#> 1 KY_Jefferson TRUE
#> 2 TN_Davidson FALSE
#> 3 OH_Hamilton FALSE
#> 4 OTHER_REGION FALSE
#> 5 KY_UNKNOWN FALSE
#> 6 <NA> FALSETwo Functions, Two Philosophies
sysPrep provides two geographic attribution functions
that differ fundamentally in which rows are
modified:
assign_treating_geography() |
assign_facility_geography() |
|
|---|---|---|
| Default output column |
region_hybrid (new column; Region
untouched) |
region_facility (new column; Region
untouched) |
| Rows reassigned | Out-of-state + OTHER_REGION + NA only |
All rows |
| In-state patient geography | Reflected as-is in region_hybrid
|
Overwritten with facility geography |
| Research question | Where did this care take place? (for non-residents only) | Where did all care take place? |
| Resulting column meaning | Mixed: residential for in-state, treating for out-of-state | Treating facility county for everyone |
Both default to writing a new column rather than overwriting
Region in place; Region itself is never
touched unless you explicitly opt into that with
overwrite = TRUE (covered at the end of this vignette).
This means calling both functions back to back gives you all three
geographies side by side with zero extra arguments:
geo_all <- essence_raw |>
dedupe(order_by = Arrived_Date_Time, keep = "last") |>
filter_care_setting(
fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
) |>
assign_treating_geography() |>
assign_facility_geography()
#> The following `FacilityType` values are not in `keep_types` and will be excluded:
#> - Primary Care
#> - Medical Specialty
#> 27 of 160 visits (16.9%) identified as out-of-state or OTHER_REGION and assigned treating facility geography in `region_hybrid`/`zip_code_hybrid`.
#> Facility geography applied to all 160 visits in `region_facility`/`zip_code_facility`.
geo_all |>
dplyr::select(hospital_name, region, region_hybrid, region_facility) |>
head(10)
#> # A tibble: 10 × 4
#> hospital_name region region_hybrid region_facility
#> <chr> <chr> <chr> <chr>
#> 1 Central Medical Center TN_Davidson KY_Jefferson KY_Jefferson
#> 2 Central Medical Center KY_Kenton KY_Kenton KY_Jefferson
#> 3 Central Medical Center KY_Campbell KY_Campbell KY_Jefferson
#> 4 Central Medical Center OH_Hamilton KY_Jefferson KY_Jefferson
#> 5 Central Medical Center OH_Franklin KY_Jefferson KY_Jefferson
#> 6 Central Medical Center KY_McCracken KY_McCracken KY_Jefferson
#> 7 Central Medical Center KY_Daviess KY_Daviess KY_Jefferson
#> 8 Central Medical Center KY_Daviess KY_Daviess KY_Jefferson
#> 9 Central Medical Center NA KY_Jefferson KY_Jefferson
#> 10 Central Medical Center KY_Campbell KY_Campbell KY_Jeffersonregion is the untouched original for every row;
region_hybrid matches it for in-state visits and
substitutes treating geography for the rest;
region_facility substitutes treating geography for
every row, including the in-state majority. One pass
over the data, three usable geographies, no
preserve_original_geographies, no manual copying.
This distinction matters for how the resulting data should be interpreted and which analyses it supports.
assign_treating_geography(): Selective Hybrid
Reassignment
assign_treating_geography() is a targeted
intervention on problem rows. It identifies visits where the
patient geography is unavailable or out-of-area and writes the treating
facility’s county, for those visits only, to region_hybrid.
region is left exactly as received from ESSENCE for every
row, in-state and out-of-state alike.
region_hybrid is a hybrid column: it
holds residential geography for patients who provided an in-state
address, and treating facility geography for patients who could not or
did not. This hybrid is appropriate for area-based incidence analyses
where you want to count as many visits as possible within your
surveillance area, and where the distinction between resident and
non-resident burden is not the primary analytical question.
# Start from deduplicated, filtered data
ed_clean <- essence_raw |>
dedupe(order_by = Arrived_Date_Time, keep = "last") |>
filter_care_setting(
fix_facility_type_vector = c("Hillside FSED", "Downtown Emergency Services")
)
#> The following `FacilityType` values are not in `keep_types` and will be excluded:
#> - Primary Care
#> - Medical Specialty
# Selective reassignment: only out-of-state and OTHER_REGION rows get a
# reassigned value in region_hybrid; region itself is never touched
treating_geo <- ed_clean |>
assign_treating_geography(site = "KY")
#> 27 of 160 visits (16.9%) identified as out-of-state or OTHER_REGION and
#> assigned treating facility geography in `region_hybrid`/`zip_code_hybrid`.
# How many visits were reassigned?
dplyr::count(treating_geo, .out_of_state)
#> # A tibble: 2 × 2
#> .out_of_state n
#> <lgl> <int>
#> 1 FALSE 133
#> 2 TRUE 27
# Compare original region against the hybrid result for reassigned rows;
# no preserve_original_geographies needed, since region was never modified
treating_geo |>
dplyr::filter(.out_of_state) |>
dplyr::select(hospital_name, region, region_hybrid) |>
head(10)
#> # A tibble: 10 × 3
#> hospital_name region region_hybrid
#> <chr> <chr> <chr>
#> 1 Central Medical Center TN_Davidson KY_Jefferson
#> 2 Central Medical Center OH_Hamilton KY_Jefferson
#> 3 Central Medical Center OH_Franklin KY_Jefferson
#> 4 Central Medical Center NA KY_Jefferson
#> 5 Central Medical Center OTHER_REGION KY_Jefferson
#> 6 Central Medical Center TN_Shelby KY_Jefferson
#> 7 Central Medical Center OH_Hamilton KY_Jefferson
#> 8 North County Hospital TN_Shelby KY_Kenton
#> 9 North County Hospital WV_Cabell KY_Kenton
#> 10 North County Hospital NA KY_KentonThe .out_of_state column (logical) marks every row where
region_hybrid differs from region. Rows where
.out_of_state = FALSE have
region_hybrid == region: the patient’s residential county,
unchanged.
If you’d rather overwrite region/zip_code
in place instead of getting a
region_hybrid/zip_code_hybrid column
(reproducing the only behavior this function had before
new_region_col/new_zip_col existed), pass
their own names explicitly and opt in with
overwrite = TRUE:
# Opt-in in-place overwrite, with an audit trail of what was replaced
treating_overwrite <- ed_clean |>
assign_treating_geography(
site = "KY",
new_region_col = "region",
new_zip_col = "zip_code",
overwrite = TRUE,
preserve_original_geographies = TRUE
)
#> 27 of 160 visits (16.9%) identified as out-of-state or OTHER_REGION and
#> assigned treating facility geography in `region`/`zip_code`.
treating_overwrite |>
dplyr::filter(.out_of_state) |>
dplyr::select(hospital_name, original_region, region) |>
head(10)
#> # A tibble: 10 × 3
#> hospital_name original_region region
#> <chr> <chr> <chr>
#> 1 Central Medical Center TN_Davidson KY_Jefferson
#> 2 Central Medical Center OH_Hamilton KY_Jefferson
#> 3 Central Medical Center OH_Franklin KY_Jefferson
#> 4 Central Medical Center NA KY_Jefferson
#> 5 Central Medical Center OTHER_REGION KY_Jefferson
#> 6 Central Medical Center TN_Shelby KY_Jefferson
#> 7 Central Medical Center OH_Hamilton KY_Jefferson
#> 8 North County Hospital TN_Shelby KY_Kenton
#> 9 North County Hospital WV_Cabell KY_Kenton
#> 10 North County Hospital NA KY_Kentonpreserve_original_geographies only has an effect in this
overwrite = TRUE mode: when it applies,
original_region/original_zip_code are added
before overwriting, retaining pre-reassignment values as an audit trail.
For in-state visits, original_region is NA
because no change was made. In the default (new-column) mode above, this
parameter is unnecessary and has no effect, since region
was never modified in the first place.
When to use
assign_treating_geography():
- Incidence estimation where you want residential geography for most patients and facility geography only as a fallback for patients without valid residential addresses.
- Trend analyses where you need to retain border-facility visits without attributing all visits to facility location.
- Sensitivity analyses: compare results with and without
OTHER_REGIONvisits using the.out_of_stateflag to filter.
assign_facility_geography(): Universal Full
Reassignment
assign_facility_geography() takes a fundamentally
different approach: every row gets a facility-geography
value in region_facility, regardless of whether the patient
was in-state, out-of-state, or unknown. region itself is
untouched, exactly as with assign_treating_geography().
region_facility never reflects where patients live; it
always reflects where they received care.
# Universal reassignment: ALL rows get a facility geography value
facility_geo <- ed_clean |> assign_facility_geography()
#> Facility geography applied to all 160 visits in
#> `region_facility`/`zip_code_facility`.
# region_facility vs. the untouched original, for every row
facility_geo |>
dplyr::select(hospital_name, hospital_region, region, region_facility) |>
head(10)
#> # A tibble: 10 × 4
#> hospital_name hospital_region region region_facility
#> <chr> <chr> <chr> <chr>
#> 1 Central Medical Center KY_Jefferson TN_Davidson KY_Jefferson
#> 2 Central Medical Center KY_Jefferson KY_Kenton KY_Jefferson
#> 3 Central Medical Center KY_Jefferson KY_Campbell KY_Jefferson
#> 4 Central Medical Center KY_Jefferson OH_Hamilton KY_Jefferson
#> 5 Central Medical Center KY_Jefferson OH_Franklin KY_Jefferson
#> 6 Central Medical Center KY_Jefferson KY_McCracken KY_Jefferson
#> 7 Central Medical Center KY_Jefferson KY_Daviess KY_Jefferson
#> 8 Central Medical Center KY_Jefferson KY_Daviess KY_Jefferson
#> 9 Central Medical Center KY_Jefferson NA KY_Jefferson
#> 10 Central Medical Center KY_Jefferson KY_Campbell KY_Jefferson
# Confirm: region_facility == hospital_region for every row
all(facility_geo$region_facility == facility_geo$hospital_region)
#> [1] TRUEUnlike assign_treating_geography(), where only
out-of-state/unknown rows differ between region and
region_hybrid, here every row differs
between region and region_facility; that’s the
“universal” part.
Is this only renaming
HospitalRegiontoRegion? Mechanically, close:assign_facility_geography()writesHospitalRegionand/orHospitalZiptoregion_facility/zip_code_facilityfor all rows, depending on which columns are present in the pull. Not all analysts include both; zip-based geography carries its own methodological considerations around spatial resolution and boundary alignment that may not be appropriate for every analysis. When onlyHospitalRegionis present, only region is reassigned. Some practitioners have applied this rename informally as an ad hoc pipeline step.sysPrepformalizes that practice: documents the intent, handles whichever geography types are available, and writes to a new column by default so the original values are always available for audit or comparison without any extra argument.assign_treating_geography(), the hybrid selective approach, has no simple column-rename equivalent, which is what motivated formalizing both methods in the same package.
Hospital-level spatial clustering (point-based) concentrates
observations at individual facility locations and can identify which
specific facilities are experiencing elevated burden, including spikes
that would be diluted or invisible at the regional level. However,
hospital-level scan statistics generally do not explicitly model
residual spatial autocorrelation or facility-specific reporting patterns
as nuisance structure. In prospective surveillance, scanning windows may
therefore interact with hospital density, catchment patterns, and data
submission variability in ways that can produce apparent clusters
independent of true changes in underlying incidence. This concern is
more pronounced in urban areas, where hospitals are denser and
submission patterns may vary across nearby facilities. Reassigning all
visits to HospitalRegion before aggregation reduces
facility-level noise by aggregating across treating facilities within a
region, allowing spacetime permutation methods and other scan statistics
to treat a rural county with one hospital comparably to an urban county
with many. Region-level binning does not eliminate spatial
autocorrelation concerns, but it reduces the influence of
facility-specific variability that is more common in high-density
areas.
A second advantage is logistical: hospital-level spatial clustering
requires latitude/longitude coordinates from the NSSP Master Facility
Table, an additional lookup outside the standard ESSENCE data pull.
Region-level attribution achieves a close approximation of
hospital-level cluster detection using HospitalRegion, a
field returned without any coordinate lookup, provided it’s
included in the pull: HospitalRegion isn’t
guaranteed on every ESSENCE pull by default, only when it’s part of your
site’s default returned columns or explicitly requested via the API’s
&field= parameter. The same logistical advantage also
supports a distinct, simpler use case: total treated volume by urban
vs. rural county, with no interest in which specific hospital saw each
visit at all, not an approximation of hospital-level detection, just
county-level counting using a field you already requested rather than a
coordinate lookup you’d have to add.
What does a county with no ESSENCE-participating hospital show?
0, always: not a small or uncertain number, a hard zero, regardless of how much true burden that county’s residents actually generate. This is a direct consequence of ESSENCE being facility-based:region_facilitycan only ever reflect counties containing a facility that submits to ESSENCE. A county with real burden but no local ED will never contribute a count under this method, no matter how the underlying incidence changes. This is worth stating plainly rather than discovering it from a suspiciously flat map.
The two approaches are complementary: hospital-based clustering offers facility-level resolution; region-based clustering offers geographic equity across areas of differing hospital density. The choice depends on whether the surveillance question is “which facility?” or “which area?”.
When to use
assign_facility_geography():
- Analyses framed around treating facility location rather than patient residence: “How many cases were treated in Jefferson County?” rather than “How many residents of Jefferson County were treated?”
- Facility service area analyses where you want all visits attributed to the county where the facility is located.
- Spatial cluster detection scoped to treating location (e.g., identifying counties with high overdose treatment burden regardless of patient origin).
- Cross-state facility comparisons where residential geography is an inappropriate common denominator.
- Comparing total treated volume across urban and rural counties
without differentiating which hospital saw each visit, useful when a
facility-to-county lookup table isn’t available but
HospitalRegion/HospitalZipalready are.
The Core Distinction: What Do the Output Columns Mean?
The choice between the two functions is a choice about what geography your analysis should be scoped to:
region_hybrid (from
assign_treating_geography()): residential geography for
in-state patients + treating facility geography for out-of-state and
unknown-residence patients.
This is a hybrid with mixed semantics. It is most useful when you want to maximize the number of visits attributed to your surveillance area while preserving the residential geography of the majority.
region_facility (from
assign_facility_geography()): treating facility geography
for all patients.
This is semantically uniform:
region_facility means the same thing for every row. It is
most useful when the analytical unit is where care was delivered, not
where patients live.
region itself, meanwhile, is never touched by either
function unless you explicitly opt into overwrite = TRUE,
so it’s always available as the untouched original, regardless of which
(or both) of these you compute. Neither function is universally correct.
The choice should be documented in your analysis methods and driven by
the research question.
Decision Guide
| Research question | Function | Rationale |
|---|---|---|
| County-level incidence rate (retain most OOS visits) | assign_treating_geography() |
Preserves residential geography for in-state majority; uses facility geography as fallback |
| Which counties have highest treatment burden? | assign_facility_geography() |
All visits attributed to treating county uniformly |
| Where do patients live who are treated at this facility? | Neither; use Region directly |
Residential geography is the question; don’t reassign |
| Spatial cluster detection by facility county | assign_facility_geography() |
Uniform attribution aligns with facility-based cluster geometry |
| Sensitivity analysis: effect of OOS exclusion | assign_treating_geography() |
Use .out_of_state flag to toggle OOS visits in/out |
| Trend analysis, mixed state border facility | assign_treating_geography() |
Avoids flattening in-state residential variation |
Output Columns: New by Default, Overwrite as an Opt-In
Every example above wrote to a new column
(region_hybrid/region_facility) and left
region/zip_code alone. That’s the default for
a reason: it’s non-destructive, so calling either function is always
safe to try, safe to re-run, and safe to chain: there’s no risk of
losing the original values just by calling the function.
new_region_col/new_zip_col control the
target column and default to
"region_hybrid"/"zip_code_hybrid"
(assign_treating_geography()) or
"region_facility"/"zip_code_facility"
(assign_facility_geography()). Pass NULL to
either to skip that geography type entirely, e.g.
new_zip_col = NULL to reassign region only.
If the target you name already exists as a column, whether that’s
region_col/zip_col itself, or any other
existing column, the function aborts rather than silently overwriting
it, unless you set overwrite = TRUE. This is what makes
reproducing the original in-place-overwrite behavior an explicit,
visible choice rather than the default:
essence_clean |>
assign_treating_geography(
new_region_col = "region",
new_zip_col = "zip_code",
overwrite = TRUE
)preserve_original_geographies = TRUE only has an effect
in this overwrite = TRUE mode: it adds
original_region/original_zip_code columns
holding the pre-overwrite values before they’re replaced. In the default
new-column mode, it has no effect and is unnecessary, since
region was never modified in the first place; it already
is the original.
For assign_treating_geography() in overwrite mode,
original_region/ original_zip_code contain
pre-reassignment values only for rows where geography changed
(.out_of_state = TRUE); for in-state visits they are
NA. For assign_facility_geography() in
overwrite mode, they contain pre-reassignment values for every
row, because every row is reassigned.