Skip to contents

Computes the five rates in cacs_acs_default_rates, such as the poverty rate, for each site and drive-time pair in the output of cacs_propagate_moe(), each with a margin of error (MOE). A margin of error is the half-width of a confidence interval. The rates are added after the input rows, as five new rows for each pair.

Usage

cacs_derive_rates(
  weighted_acs,
  rates = cacs_acs_default_rates,
  formula_dispatch = "general_ratio_conservative",
  verbose = TRUE,
  ...
)

Arguments

weighted_acs

A tibble returned by cacs_propagate_moe(), from which rows may have been removed (for example with [ or dplyr::filter()) but whose columns and attributes are kept, including cacs_aggregation_carriers. A tibble returned by cacs_intersect_weight() can also be used; the margins of error of the rates are then at the 90 percent level. The rates are computed from cacs_aggregation_carriers for each site and drive-time pair that still has a row, so removing the rows of a rate's codes does not change the rate.

rates

A named list giving, for each rate, the American Community Survey (ACS) codes of the numerator and the denominator. Only cacs_acs_default_rates (the default), the list of the five built-in rates, is accepted; any other list, including a subset or a reordering of it, gives an error.

formula_dispatch

A string choosing the margin-of-error formula for the rates: "general_ratio_conservative" (the default), "auto", or "proportion_subset". The last two give the same result: the proportion formula for poverty_rate and labor_force_participation only, and the ratio formula for the other three (see Details).

verbose

A logical value. With TRUE (the default), the function shows a one-line summary when the step finishes (see the "Progress messages" section of cacs_run()).

...

Not used: any argument given here, such as level = 0.95, gives an error. The confidence level of the rates is set by the level argument of cacs_propagate_moe().

Value

The tibble weighted_acs with its rows unchanged and in the same order, followed by the rate rows. On the rate rows:

variable, estimand_family

The name of the rate, such as "poverty_rate", and "derived_rate".

estimate, moe

The rate (not a percentage) and its margin of error.

moe_formula_requested, moe_formula_effective

The formula chosen for the rate (see the table in Details) and the formula used, which differs only where the ratio formula replaced the proportion formula. With formula_dispatch = "proportion_subset", the three rates that keep the ratio formula have "general_ratio_conservative" in both columns.

moe_fallback, moe_fallback_reason

TRUE with "negative_variance" or "zero_denominator" when the chosen formula could not be used (see Details and the "Rates that are NA" section), and FALSE with "n/a" otherwise, including when failure_origin is "carrier".

weight_sum

The smaller of the sums of the coverage weights for the numerator and for the denominator (a tract's coverage weight is the share of its area inside the drive-time area); NA when failure_origin is "carrier".

n_tracts, n_tracts_num, n_tracts_den

See the "Tract counts for rates" section.

failure_origin

"none", or "carrier" when a count or margin of error that the rate needs is missing (see the "Rates that are NA" section).

The other columns are described in the Value section of cacs_run().

The attributes of weighted_acs are kept, except cacs_aggregation_carriers, and cacs_confidence_level is set to the confidence level used. Two attributes are added:

cacs_rate_provenance

A list recording how the rates were computed. It gives the rates and the formula chosen for each, the value of formula_dispatch, and the rates that "proportion_subset" left on the ratio formula. It also counts the rate rows: all of them, and those with a substituted formula, a zero denominator, failure_origin = "carrier", or a rate outside the check ranges. The last fields are the confidence level and its z value, the time the rates were computed, and the package version.

cacs_rate_audit

The result of the range check described in the "Checking rates against ranges" section.

Details

Each rate is the ratio of two weighted counts of ACS estimates. The two counts and their variances are read from the cacs_aggregation_carriers attribute that cacs_intersect_weight() attaches. The margin of error is computed from them at the confidence level that cacs_propagate_moe() records in the cacs_confidence_level attribute (the 90 percent level if there is none). The result does not keep cacs_aggregation_carriers, so cacs_derive_rates() gives an error if it is run on its own result.

The numerator of each of the five rates is part of its denominator, so the rates are proportions, for which the Census Bureau's handbook gives the proportion formula, C1 (U.S. Census Bureau 2020, chapter 8). By default, the margin of error of every rate is computed with the ratio formula, C2, which for the same rate is never narrower than the margin the proportion formula gives. formula_dispatch can choose the proportion formula for two of the five rates. The table gives the formula used for each rate with each value of formula_dispatch.

RateDefault"auto""proportion_subset"
poverty_rateratioproportionproportion
labor_force_participationratioproportionproportion
snap_rateratioratioratio
ssi_rateratioratioratio
unemp_rateratioratioratio

"auto" and "proportion_subset" give the same rates and margins of error. "proportion_subset" also gives a warning naming the three rates that keep the ratio formula, and lists them in the formula_downgraded_rates element of the cacs_rate_provenance attribute, a list recording how the rates were computed. The warning calls those three rates ineligible for the proportion formula. That wording describes the package's list of rates and not the data: the numerator of each of the three is part of its denominator. Both formulas are given in the "MOE formula families" section of cacs_propagate_moe().

Under "auto" and "proportion_subset", the three rates in the table keep the ratio formula whatever the values are; the choice is built into the package and does not depend on the data. unemp_rate keeps it to reproduce the 2025 analysis the package was first written for, and the package gives no reason for snap_rate and ssi_rate. vignette("theory-derived-rates", package = "catchmentACS") describes how the proportion formula can be computed for them from the estimate and moe of the rows of their two ACS codes.

When the value under the square root of the proportion formula is negative, the row uses the ratio formula instead (see cacs_propagate_moe()), with moe_fallback = TRUE and moe_fallback_reason = "negative_variance".

Rates that are NA

A rate and its margin of error are NA when the numerator or the denominator, or the margin of error of either, is missing for the pair. This happens when a code of the rate has no rows in the ACS data, when a tract in the area has a missing estimate or margin of error for one of the two codes (including a Census Bureau code that cacs_intersect_weight() sets to NA), and when no tract is left for the pair. The codes of all five rates are in cacs_acs_default_vars, the default variables of cacs_acs_prefetch() and cacs_run(). Rates that are NA for these reasons have "carrier" in failure_origin, the column that records the step at which a row failed; the value is named after the cacs_aggregation_carriers attribute, which holds the weighted totals. A rate whose denominator is zero is also NA, as is its margin of error, with "zero_denominator" in moe_fallback_reason and TRUE in moe_fallback.

One warning at the end counts these rows and the rows where the ratio formula replaced the proportion formula. Its class is catchmentACS_warning_runtime; when it counts only rows with failure_origin = "carrier", it also has the class catchmentACS_warning_carrier_missing.

Tract counts for rates

On rate rows, n_tracts_num and n_tracts_den give the numbers of tracts combined for the numerator and for the denominator (the values of n_tracts on the rows of those two ACS codes), and n_tracts is NA. On the rows of the ACS variables, n_tracts_num and n_tracts_den are NA, and they are also NA on rate rows with failure_origin = "carrier".

When, in the ACS data, a tract in the area has a row for one of the two codes of a rate and not for the other, for example after rows with a missing estimate have been removed, the numerator and the denominator are totals over different sets of tracts. The rate is then no longer the share of one population: it can be much larger or smaller than the share for the area, and even above 1, and no warning is given. The two counts then differ unless the two codes lack the same number of tracts, so equal counts do not rule this out. Rows whose estimate is missing, kept as NA rows rather than removed, make the rate NA, with a warning (see the "Rates that are NA" section). cacs_acs_validate() shows how to check that every tract has one row for each variable.

Checking rates against ranges

With options(catchmentACS.audit_rates = TRUE), each rate is compared with a fixed range: 0 to 0.6 for poverty_rate, 0 to 0.5 for snap_rate, 0 to 0.15 for ssi_rate, 0 to 0.3 for unemp_rate, and 0.3 to 0.85 for labor_force_participation. For each rate with values outside its range, a warning of class catchmentACS_warning_rate_out_of_range gives the number of those rows and the smallest and largest of their values; the rates are not changed. The cacs_rate_audit attribute records whether the check was on (enabled), the number of rate rows outside the ranges (n_out_of_bounds), and those rows (rows). It is attached even when the check is off (the default), with enabled = FALSE and no rows.

References

U.S. Census Bureau (2020). Understanding and Using American Community Survey Data: What All Data Users Need to Know. Chapter 8, Calculating Measures of Error for Derived Estimates.

Examples

# Turn the cache off while this example runs (see ?cacs_set_cache).
old <- options(catchmentACS.cache_enabled = FALSE)

# Example data bundled with the package: the drive-time areas are circles
# with a radius of 1 km per minute, and the ACS data are made up.
library(sf)
iso <- readRDS(system.file("extdata", "legacy_2025_isochrones.rds",
                           package = "catchmentACS"))
acs <- readRDS(system.file("extdata", "sample_alabama_subset.rds",
                           package = "catchmentACS"))
iso_07 <- iso[iso$site_id == "AL_SITE_07" & iso$drive_time_min == 10, ]
agg <- cacs_intersect_weight(iso_sf = iso_07, acs_sf = acs,
                             verbose = FALSE)
prop <- cacs_propagate_moe(agg, verbose = FALSE)

# The five rates, with margins of error at the 90 percent level
out <- cacs_derive_rates(prop, verbose = FALSE)
is_rate <- out$estimand_family == "derived_rate"
out[is_rate, c("variable", "estimate", "moe", "moe_formula_effective")]
#> # A tibble: 5 × 4
#>   variable                  estimate     moe moe_formula_effective     
#>   <chr>                        <dbl>   <dbl> <chr>                     
#> 1 poverty_rate                0.602  0.0984  general_ratio_conservative
#> 2 snap_rate                   0.0426 0.00557 general_ratio_conservative
#> 3 ssi_rate                    0.165  0.0360  general_ratio_conservative
#> 4 unemp_rate                  0.0630 0.0149  general_ratio_conservative
#> 5 labor_force_participation   0.846  0.112   general_ratio_conservative

# Margins of error with the default and with formula_dispatch = "auto"
out_auto <- cacs_derive_rates(prop, formula_dispatch = "auto",
                              verbose = FALSE)
tibble::tibble(variable = out$variable[is_rate],
               moe_default = out$moe[is_rate],
               moe_auto = out_auto$moe[is_rate],
               formula_auto = out_auto$moe_formula_effective[is_rate])
#> # A tibble: 5 × 4
#>   variable                  moe_default moe_auto formula_auto              
#>   <chr>                           <dbl>    <dbl> <chr>                     
#> 1 poverty_rate                  0.0984   0.0858  proportion_subset         
#> 2 snap_rate                     0.00557  0.00557 general_ratio_conservative
#> 3 ssi_rate                      0.0360   0.0360  general_ratio_conservative
#> 4 unemp_rate                    0.0149   0.0149  general_ratio_conservative
#> 5 labor_force_participation     0.112    0.0904  proportion_subset         

options(old)