Skip to contents

Computes a margin of error (MOE) for each estimate in the output of cacs_intersect_weight(), at the confidence level given by level. A margin of error is the half-width of a confidence interval. The formula depends on the kind of estimate and is recorded in the moe_formula_effective column. American Community Survey (ACS) margins of error are published at the 90 percent level; at the default level = 0.9, the margins of error of counts, medians, and per-person values are the same as those that cacs_intersect_weight() returns.

Usage

cacs_propagate_moe(data, formula = NULL, level = 0.9, verbose = TRUE, ...)

Arguments

data

A tibble returned by cacs_intersect_weight(), from which rows may have been removed (for example with [ or dplyr::filter()) but whose columns and attributes are kept, including cacs_aggregation_carriers (a table of the weighted totals and means with their variances, which this function reads). Rows for ACS codes that cacs_intersect_weight() does not combine (estimand_family is "metadata_only"), such as codes of tables whose names begin with C or S, make the function stop with an error; it runs once those rows are removed. The row kept for a site and drive-time pair with no tracts is returned unchanged.

formula

A string, a list of strings named by values of variable, or NULL (the default), giving the formula for each row: "weighted_sum", "weighted_mean", "proportion_subset", or "general_ratio_conservative" (see the "MOE formula families" section). NULL gives each row the formula for its kind of estimate, and a string is used for every row. In a list, the rows not named keep the formula they would get with NULL. Each row accepts only the formula for its kind of estimate; any other formula, a name not found in variable, or an unknown formula gives an error.

level

A single number between 0 and 1 (not a percentage) giving the confidence level of the margins of error. The default is 0.9. Any other value, such as 0, 1, 90, NA, "0.9", or c(0.9, 0.95), gives an error.

verbose

A logical value. With TRUE (the default), the function shows a one-line summary when the step finishes (see the "Progress messages" section of cacs_run()).

...

Not used. The name fallback_chain_max is accepted with a warning and has no effect; any other argument gives an error.

Value

The tibble data with the same rows in the same order, in which moe (the margin of error) is recomputed at level. The columns that record how it was computed are rewritten: moe_formula_requested and moe_formula_effective (the formula chosen and the one used), and moe_fallback and moe_fallback_reason (whether the chosen formula could not be used, and why). Count, median, and per-person rows keep moe_fallback = FALSE even when their margin of error is NA. Four columns are added from the cacs_aggregation_carriers attribute: est_total and var_total_raw, the weighted sum of the tract estimates and its variance (used for counts), and est_mean and var_mean_raw, the weighted average and its variance (used for medians and per-person values). With options(cacs.return_se = TRUE), an se column is added too: the standard error, moe divided by 1.645 at the 90 percent level or by the normal quantile for another level.

The attributes of data are kept, and two are added. cacs_confidence_level is the value of level, which cacs_derive_rates() uses for the rates. cacs_moe_provenance is a list recording the confidence level, the formula argument, and the numbers of rows computed with Families A and B (see the "MOE formula families" section). Its counts for rates are always 0, because the rates are added later by cacs_derive_rates(), which counts them in its cacs_rate_provenance attribute.

Details

The margins of error are combined as if the census tract estimates were independent. The Census Bureau's handbook for ACS data users notes that its approximation formulas leave out the covariance between estimates, so a margin of error can be too small or too large depending on the correlation between them (U.S. Census Bureau 2020, chapter 8). The margin of error of a count, a median, or a per-person value is too small if the tract estimates are positively correlated. For a rate, errors that move the numerator and the denominator in the same direction partly offset each other in the ratio, and the package does not compute the net effect for a given rate. The area-based weights of the tracts are treated as fixed, so the margins of error leave out two more sources of error. One is the assumption that whatever a variable counts is spread evenly over each tract's area. The other, for medians and per-person values, is the weighting of tracts by area.

MOE formula families

By default, each row gets the formula for the kind of estimate in its estimand_family column. The four formulas are called Families A, B, C1, and C2; the letters also appear in warnings and in the row counts of the cacs_moe_provenance attribute. The formulas give the margin of error at the 90 percent level, the level of the tract margins of error \(M_i\); at another level, the result is multiplied by \(z / 1.645\), where \(z\) is the normal quantile for the level, qnorm(1 - (1 - level) / 2).

  • Family A, recorded as "weighted_sum", is used for counts ("spatial_total"). The margin of error is \(\sqrt{\sum_i (w_i M_i)^2}\), where \(w_i\) is the coverage weight of tract \(i\), the share of its area inside the drive-time area. This is the handbook's formula for a sum, applied to the weighted tract estimates.

  • Family B, recorded as "weighted_mean", is used for medians and per-person values ("median_proxy" and "area_weighted_scalar_proxy"). It is the same formula, with \(w_i\) proportional to the area that tract \(i\) shares with the drive-time area and summing to one. This is an approximation made by the package rather than one given in the handbook.

  • Family C1, recorded as "proportion_subset", is the handbook's formula for a proportion, a ratio whose numerator is part of its denominator. For the rate \(\hat p = \hat X_{\mathrm{num}} / \hat X_{\mathrm{den}}\) of two weighted counts with Family A margins of error \(M_{\mathrm{num}}\) and \(M_{\mathrm{den}}\), the margin of error is \(\sqrt{M_{\mathrm{num}}^2 - \hat p^2 M_{\mathrm{den}}^2} / \hat X_{\mathrm{den}}\).

  • Family C2, recorded as "general_ratio_conservative", is the handbook's formula for a ratio whose numerator is not part of its denominator: \(\sqrt{M_{\mathrm{num}}^2 + \hat p^2 M_{\mathrm{den}}^2} / \left| \hat X_{\mathrm{den}} \right|\). The package divides by the absolute value of the denominator; for the five rates, whose denominators are counts, that is the denominator itself. Like the other formulas it leaves out the correlation between tract estimates (see Details).

C1 and C2 are used for the rows of the five rates in cacs_acs_default_rates ("derived_rate"). By default every rate uses C2, the ratio formula. The numerator of each of the five rates is part of its denominator, so the rates are proportions, for which the handbook gives C1, the proportion formula. The two formulas differ by \(2 \hat p^2 M_{\mathrm{den}}^2\) under the square root, so C2 never gives the narrower margin; how much wider it is depends on \(\hat p M_{\mathrm{den}}\) relative to \(M_{\mathrm{num}}\) and varies from rate to rate.

If the value under the square root of C1 is negative, the row uses C2 instead, as the handbook advises (U.S. Census Bureau 2020, chapter 8). The substitution is recorded in the row: "general_ratio_conservative" in moe_formula_effective, TRUE in moe_fallback, and "negative_variance" in moe_fallback_reason. How rates that cannot be computed are marked is described in the "Rates that are NA" section of cacs_derive_rates().

The output of cacs_intersect_weight() has no rate rows, so this function applies only Families A and B; cacs_derive_rates() applies C1 and C2 to the rates, as chosen by its formula_dispatch argument, at the level set here.

References

U.S. Census Bureau (2020). Understanding and Using American Community Survey Data: What All Data Users Need to Know. Chapter 8, Calculating Measures of Error for Derived Estimates.

Examples

# Turn the cache off while this example runs (see ?cacs_set_cache).
old <- options(catchmentACS.cache_enabled = FALSE)

# Example data bundled with the package: the drive-time areas are circles
# with a radius of 1 km per minute, and the ACS data are made up.
library(sf)
iso <- readRDS(system.file("extdata", "legacy_2025_isochrones.rds",
                           package = "catchmentACS"))
acs <- readRDS(system.file("extdata", "sample_alabama_subset.rds",
                           package = "catchmentACS"))
iso_07 <- iso[iso$site_id == "AL_SITE_07" & iso$drive_time_min == 10, ]
agg <- cacs_intersect_weight(iso_sf = iso_07, acs_sf = acs,
                             verbose = FALSE)

# Margins of error at the default 90 percent level and at 95 percent
prop <- cacs_propagate_moe(agg, verbose = FALSE)
prop_95 <- cacs_propagate_moe(agg, level = 0.95, verbose = FALSE)
tibble::tibble(variable = prop$variable, estimate = prop$estimate,
               moe_90 = prop$moe, moe_95 = prop_95$moe,
               formula = prop$moe_formula_effective)
#> # A tibble: 14 × 5
#>    variable   estimate   moe_90   moe_95 formula      
#>    <chr>         <dbl>    <dbl>    <dbl> <chr>        
#>  1 B01003_001   4643.    612.     729.   weighted_sum 
#>  2 B11001_001   1557.    309.     368.   weighted_sum 
#>  3 B17001_001   1961.    111.     132.   weighted_sum 
#>  4 B17001_002   1181.    181.     216.   weighted_sum 
#>  5 B19013_001  68037.  16160.   19255.   weighted_mean
#>  6 B19056_001   1033.     76.1     90.7  weighted_sum 
#>  7 B19056_002    171.     35.0     41.7  weighted_sum 
#>  8 B19301_001  16725.   2356.    2807.   weighted_mean
#>  9 B22003_001   1746.    158.     188.   weighted_sum 
#> 10 B22003_002     74.5     7.02     8.36 weighted_sum 
#> 11 B23025_001   2090.    114.     136.   weighted_sum 
#> 12 B23025_002   1767.    212.     253.   weighted_sum 
#> 13 B23025_003   2185.    261.     311.   weighted_sum 
#> 14 B23025_005    138.     28.0     33.4  weighted_sum 

options(old)