Compute margins of error for drive-time area estimates
Source:R/moe-helpers.R
cacs_propagate_moe.RdComputes a margin of error (MOE) for each estimate in the output of
cacs_intersect_weight(), at the confidence level given by level. A
margin of error is the half-width of a confidence interval. The formula
depends on the kind of estimate and is recorded in the
moe_formula_effective column. American Community Survey (ACS) margins
of error are published at the 90 percent level; at the default
level = 0.9, the margins of error of counts, medians, and per-person
values are the same as those that cacs_intersect_weight() returns.
Arguments
- data
A tibble returned by
cacs_intersect_weight(), from which rows may have been removed (for example with[ordplyr::filter()) but whose columns and attributes are kept, includingcacs_aggregation_carriers(a table of the weighted totals and means with their variances, which this function reads). Rows for ACS codes thatcacs_intersect_weight()does not combine (estimand_familyis"metadata_only"), such as codes of tables whose names begin withCorS, make the function stop with an error; it runs once those rows are removed. The row kept for a site and drive-time pair with no tracts is returned unchanged.- formula
A string, a list of strings named by values of
variable, orNULL(the default), giving the formula for each row:"weighted_sum","weighted_mean","proportion_subset", or"general_ratio_conservative"(see the "MOE formula families" section).NULLgives each row the formula for its kind of estimate, and a string is used for every row. In a list, the rows not named keep the formula they would get withNULL. Each row accepts only the formula for its kind of estimate; any other formula, a name not found invariable, or an unknown formula gives an error.- level
A single number between 0 and 1 (not a percentage) giving the confidence level of the margins of error. The default is
0.9. Any other value, such as0,1,90,NA,"0.9", orc(0.9, 0.95), gives an error.- verbose
A logical value. With
TRUE(the default), the function shows a one-line summary when the step finishes (see the "Progress messages" section ofcacs_run()).- ...
Not used. The name
fallback_chain_maxis accepted with a warning and has no effect; any other argument gives an error.
Value
The tibble data with the same rows in the same order, in which
moe (the margin of error) is recomputed at level. The columns that
record how it was computed are rewritten: moe_formula_requested and
moe_formula_effective (the formula chosen and the one used), and
moe_fallback and moe_fallback_reason (whether the chosen formula
could not be used, and why). Count, median, and per-person rows keep
moe_fallback = FALSE even when their margin of error is NA. Four
columns are added from the cacs_aggregation_carriers attribute:
est_total and var_total_raw, the weighted sum of the tract estimates
and its variance (used for counts), and est_mean and var_mean_raw,
the weighted average and its variance (used for medians and per-person
values). With options(cacs.return_se = TRUE), an se column is added
too: the standard error, moe divided by 1.645 at the 90 percent level
or by the normal quantile for another level.
The attributes of data are kept, and two are added.
cacs_confidence_level is the value of level, which
cacs_derive_rates() uses for the rates. cacs_moe_provenance is a
list recording the confidence level, the formula argument, and the
numbers of rows computed with Families A and B (see the "MOE formula
families" section). Its counts for rates are always 0, because the
rates are added later by cacs_derive_rates(), which counts them in its
cacs_rate_provenance attribute.
Details
The margins of error are combined as if the census tract estimates were independent. The Census Bureau's handbook for ACS data users notes that its approximation formulas leave out the covariance between estimates, so a margin of error can be too small or too large depending on the correlation between them (U.S. Census Bureau 2020, chapter 8). The margin of error of a count, a median, or a per-person value is too small if the tract estimates are positively correlated. For a rate, errors that move the numerator and the denominator in the same direction partly offset each other in the ratio, and the package does not compute the net effect for a given rate. The area-based weights of the tracts are treated as fixed, so the margins of error leave out two more sources of error. One is the assumption that whatever a variable counts is spread evenly over each tract's area. The other, for medians and per-person values, is the weighting of tracts by area.
MOE formula families
By default, each row gets the formula for the kind of estimate in its
estimand_family column. The four formulas are called Families A, B, C1,
and C2; the letters also appear in warnings and in the row counts of the
cacs_moe_provenance attribute. The formulas give the margin of error at
the 90 percent level, the level of the tract margins of error \(M_i\);
at another level, the result is multiplied by \(z / 1.645\), where
\(z\) is the normal quantile for the level,
qnorm(1 - (1 - level) / 2).
Family A, recorded as
"weighted_sum", is used for counts ("spatial_total"). The margin of error is \(\sqrt{\sum_i (w_i M_i)^2}\), where \(w_i\) is the coverage weight of tract \(i\), the share of its area inside the drive-time area. This is the handbook's formula for a sum, applied to the weighted tract estimates.Family B, recorded as
"weighted_mean", is used for medians and per-person values ("median_proxy"and"area_weighted_scalar_proxy"). It is the same formula, with \(w_i\) proportional to the area that tract \(i\) shares with the drive-time area and summing to one. This is an approximation made by the package rather than one given in the handbook.Family C1, recorded as
"proportion_subset", is the handbook's formula for a proportion, a ratio whose numerator is part of its denominator. For the rate \(\hat p = \hat X_{\mathrm{num}} / \hat X_{\mathrm{den}}\) of two weighted counts with Family A margins of error \(M_{\mathrm{num}}\) and \(M_{\mathrm{den}}\), the margin of error is \(\sqrt{M_{\mathrm{num}}^2 - \hat p^2 M_{\mathrm{den}}^2} / \hat X_{\mathrm{den}}\).Family C2, recorded as
"general_ratio_conservative", is the handbook's formula for a ratio whose numerator is not part of its denominator: \(\sqrt{M_{\mathrm{num}}^2 + \hat p^2 M_{\mathrm{den}}^2} / \left| \hat X_{\mathrm{den}} \right|\). The package divides by the absolute value of the denominator; for the five rates, whose denominators are counts, that is the denominator itself. Like the other formulas it leaves out the correlation between tract estimates (see Details).
C1 and C2 are used for the rows of the five rates in
cacs_acs_default_rates ("derived_rate"). By default every rate uses
C2, the ratio formula. The numerator of each of the five rates is part of
its denominator, so the rates are proportions, for which the handbook
gives C1, the proportion formula. The two formulas differ by
\(2 \hat p^2 M_{\mathrm{den}}^2\) under the square
root, so C2 never gives the narrower margin; how much wider it is depends
on \(\hat p M_{\mathrm{den}}\) relative to
\(M_{\mathrm{num}}\) and varies from rate to rate.
If the value under the square root of C1 is negative, the row uses C2
instead, as the handbook advises (U.S. Census Bureau 2020, chapter 8).
The substitution is recorded in the row: "general_ratio_conservative"
in moe_formula_effective, TRUE in moe_fallback, and
"negative_variance" in moe_fallback_reason. How rates that cannot be
computed are marked is described in the "Rates that are NA" section of
cacs_derive_rates().
The output of cacs_intersect_weight() has no rate rows, so this function
applies only Families A and B; cacs_derive_rates() applies C1 and C2 to
the rates, as chosen by its formula_dispatch argument, at the level set
here.
References
U.S. Census Bureau (2020). Understanding and Using American Community Survey Data: What All Data Users Need to Know. Chapter 8, Calculating Measures of Error for Derived Estimates.
See also
cacs_se_to_moe() and cacs_moe_to_se() convert between
margins of error and standard errors. The mathematics behind these
formulas is in
vignette("theory-moe-propagation", package = "catchmentACS")
(https://joonho112.github.io/catchmentACS/articles/theory-moe-propagation.html).
Other steps of the calculation:
cacs_acs_prefetch(),
cacs_derive_rates(),
cacs_intersect_weight(),
cacs_isochrone(),
cacs_run()
Examples
# Turn the cache off while this example runs (see ?cacs_set_cache).
old <- options(catchmentACS.cache_enabled = FALSE)
# Example data bundled with the package: the drive-time areas are circles
# with a radius of 1 km per minute, and the ACS data are made up.
library(sf)
iso <- readRDS(system.file("extdata", "legacy_2025_isochrones.rds",
package = "catchmentACS"))
acs <- readRDS(system.file("extdata", "sample_alabama_subset.rds",
package = "catchmentACS"))
iso_07 <- iso[iso$site_id == "AL_SITE_07" & iso$drive_time_min == 10, ]
agg <- cacs_intersect_weight(iso_sf = iso_07, acs_sf = acs,
verbose = FALSE)
# Margins of error at the default 90 percent level and at 95 percent
prop <- cacs_propagate_moe(agg, verbose = FALSE)
prop_95 <- cacs_propagate_moe(agg, level = 0.95, verbose = FALSE)
tibble::tibble(variable = prop$variable, estimate = prop$estimate,
moe_90 = prop$moe, moe_95 = prop_95$moe,
formula = prop$moe_formula_effective)
#> # A tibble: 14 × 5
#> variable estimate moe_90 moe_95 formula
#> <chr> <dbl> <dbl> <dbl> <chr>
#> 1 B01003_001 4643. 612. 729. weighted_sum
#> 2 B11001_001 1557. 309. 368. weighted_sum
#> 3 B17001_001 1961. 111. 132. weighted_sum
#> 4 B17001_002 1181. 181. 216. weighted_sum
#> 5 B19013_001 68037. 16160. 19255. weighted_mean
#> 6 B19056_001 1033. 76.1 90.7 weighted_sum
#> 7 B19056_002 171. 35.0 41.7 weighted_sum
#> 8 B19301_001 16725. 2356. 2807. weighted_mean
#> 9 B22003_001 1746. 158. 188. weighted_sum
#> 10 B22003_002 74.5 7.02 8.36 weighted_sum
#> 11 B23025_001 2090. 114. 136. weighted_sum
#> 12 B23025_002 1767. 212. 253. weighted_sum
#> 13 B23025_003 2185. 261. 311. weighted_sum
#> 14 B23025_005 138. 28.0 33.4 weighted_sum
options(old)