Compute rates and their margins of error for drive-time areas
Source:R/derive-rates.R
cacs_derive_rates.RdComputes the five rates in cacs_acs_default_rates, such as the
poverty rate, for each site and drive-time pair in the output of
cacs_propagate_moe(), each with a margin of error (MOE). A margin of
error is the half-width of a confidence interval. The rates are added
after the input rows, as five new rows for each pair.
Usage
cacs_derive_rates(
weighted_acs,
rates = cacs_acs_default_rates,
formula_dispatch = "general_ratio_conservative",
verbose = TRUE,
...
)Arguments
- weighted_acs
A tibble returned by
cacs_propagate_moe(), from which rows may have been removed (for example with[ordplyr::filter()) but whose columns and attributes are kept, includingcacs_aggregation_carriers. A tibble returned bycacs_intersect_weight()can also be used; the margins of error of the rates are then at the 90 percent level. The rates are computed fromcacs_aggregation_carriersfor each site and drive-time pair that still has a row, so removing the rows of a rate's codes does not change the rate.- rates
A named list giving, for each rate, the American Community Survey (ACS) codes of the numerator and the denominator. Only
cacs_acs_default_rates(the default), the list of the five built-in rates, is accepted; any other list, including a subset or a reordering of it, gives an error.- formula_dispatch
A string choosing the margin-of-error formula for the rates:
"general_ratio_conservative"(the default),"auto", or"proportion_subset". The last two give the same result: the proportion formula forpoverty_rateandlabor_force_participationonly, and the ratio formula for the other three (see Details).- verbose
A logical value. With
TRUE(the default), the function shows a one-line summary when the step finishes (see the "Progress messages" section ofcacs_run()).- ...
Not used: any argument given here, such as
level = 0.95, gives an error. The confidence level of the rates is set by thelevelargument ofcacs_propagate_moe().
Value
The tibble weighted_acs with its rows unchanged and in the same
order, followed by the rate rows. On the rate rows:
variable,estimand_familyThe name of the rate, such as
"poverty_rate", and"derived_rate".estimate,moeThe rate (not a percentage) and its margin of error.
moe_formula_requested,moe_formula_effectiveThe formula chosen for the rate (see the table in Details) and the formula used, which differs only where the ratio formula replaced the proportion formula. With
formula_dispatch = "proportion_subset", the three rates that keep the ratio formula have"general_ratio_conservative"in both columns.moe_fallback,moe_fallback_reasonTRUEwith"negative_variance"or"zero_denominator"when the chosen formula could not be used (see Details and the "Rates that are NA" section), andFALSEwith"n/a"otherwise, including whenfailure_originis"carrier".weight_sumThe smaller of the sums of the coverage weights for the numerator and for the denominator (a tract's coverage weight is the share of its area inside the drive-time area);
NAwhenfailure_originis"carrier".n_tracts,n_tracts_num,n_tracts_denSee the "Tract counts for rates" section.
failure_origin"none", or"carrier"when a count or margin of error that the rate needs is missing (see the "Rates that are NA" section).
The other columns are described in the Value section of cacs_run().
The attributes of weighted_acs are kept, except
cacs_aggregation_carriers, and cacs_confidence_level is set to the
confidence level used. Two attributes are added:
cacs_rate_provenanceA list recording how the rates were computed. It gives the rates and the formula chosen for each, the value of
formula_dispatch, and the rates that"proportion_subset"left on the ratio formula. It also counts the rate rows: all of them, and those with a substituted formula, a zero denominator,failure_origin = "carrier", or a rate outside the check ranges. The last fields are the confidence level and its z value, the time the rates were computed, and the package version.cacs_rate_auditThe result of the range check described in the "Checking rates against ranges" section.
Details
Each rate is the ratio of two weighted counts of ACS estimates. The two
counts and their variances are read from the cacs_aggregation_carriers
attribute that cacs_intersect_weight() attaches. The margin of error is
computed from them at the confidence level that cacs_propagate_moe()
records in the cacs_confidence_level attribute (the 90 percent level if
there is none). The result does not keep cacs_aggregation_carriers, so
cacs_derive_rates() gives an error if it is run on its own result.
The numerator of each of the five rates is part of its denominator, so
the rates are proportions, for which the Census Bureau's handbook gives
the proportion formula, C1 (U.S. Census Bureau 2020, chapter 8). By
default, the margin of error of every rate is computed with the ratio
formula, C2, which for the same rate is never narrower than the margin the
proportion formula gives. formula_dispatch can choose the proportion
formula for two of the five rates. The table gives the formula used for
each rate with each value of formula_dispatch.
| Rate | Default | "auto" | "proportion_subset" |
poverty_rate | ratio | proportion | proportion |
labor_force_participation | ratio | proportion | proportion |
snap_rate | ratio | ratio | ratio |
ssi_rate | ratio | ratio | ratio |
unemp_rate | ratio | ratio | ratio |
"auto" and "proportion_subset" give the same rates and margins of
error. "proportion_subset" also gives a warning naming the three rates
that keep the ratio formula, and lists them in the
formula_downgraded_rates element of the cacs_rate_provenance
attribute, a list recording how the rates were computed. The warning calls
those three rates ineligible for the proportion formula. That wording
describes the package's list of rates and not the data: the numerator of
each of the three is part of its denominator. Both formulas are given in
the "MOE formula families" section of cacs_propagate_moe().
Under "auto" and "proportion_subset", the three rates in the table keep
the ratio formula whatever the values are; the choice is built into the
package and does not depend on the data. unemp_rate keeps it to
reproduce the 2025 analysis the package was first written for, and the
package gives no reason for snap_rate and ssi_rate.
vignette("theory-derived-rates", package = "catchmentACS") describes how
the proportion formula can be computed for them from the estimate and
moe of the rows of their two ACS codes.
When the value under the square root of the proportion formula is
negative, the row uses the ratio formula instead (see
cacs_propagate_moe()), with moe_fallback = TRUE and
moe_fallback_reason = "negative_variance".
Rates that are NA
A rate and its margin of error are NA when the numerator or the
denominator, or the margin of error of either, is missing for the pair.
This happens when a code of the rate has no rows in the ACS data, when a
tract in the area has a missing estimate or margin of error for one of
the two codes (including a Census Bureau code that
cacs_intersect_weight() sets to NA), and when no tract is left for the
pair. The codes of all
five rates are in cacs_acs_default_vars, the default variables of
cacs_acs_prefetch() and cacs_run(). Rates that are NA for these
reasons have "carrier" in failure_origin, the column that records the
step at which a row failed; the value is named after the
cacs_aggregation_carriers attribute, which holds the weighted totals. A
rate whose denominator is zero is also NA, as is its margin of error,
with "zero_denominator" in moe_fallback_reason and TRUE in
moe_fallback.
One warning at the end counts these rows and the rows where the ratio
formula replaced the proportion formula. Its class is
catchmentACS_warning_runtime; when it counts only rows with
failure_origin = "carrier", it also has the class
catchmentACS_warning_carrier_missing.
Tract counts for rates
On rate rows, n_tracts_num and n_tracts_den give the numbers of
tracts combined for the numerator and for the denominator (the values of
n_tracts on the rows of those two ACS codes), and n_tracts is NA.
On the rows of the ACS variables, n_tracts_num and n_tracts_den are
NA, and they are also NA on rate rows with
failure_origin = "carrier".
When, in the ACS data, a tract in the area has a row for one of the two
codes of a rate and not for the other, for example after rows with a
missing estimate have been removed, the numerator and the denominator are
totals over different sets of tracts. The rate is then no longer the
share of one population: it can be much larger or smaller than the share
for the area, and even above 1, and no warning is given. The two counts
then differ unless the two codes lack the same number of tracts, so equal
counts do not rule this out. Rows whose estimate is missing, kept as NA
rows rather than removed, make the rate NA, with a warning (see the
"Rates that are NA" section). cacs_acs_validate() shows how to check
that every tract has one row for each variable.
Checking rates against ranges
With options(catchmentACS.audit_rates = TRUE), each rate is compared
with a fixed range: 0 to 0.6 for poverty_rate, 0 to 0.5 for
snap_rate, 0 to 0.15 for ssi_rate, 0 to 0.3 for unemp_rate, and
0.3 to 0.85 for labor_force_participation. For each rate with values
outside its range, a warning of class
catchmentACS_warning_rate_out_of_range gives the number of those rows
and the smallest and largest of their values; the rates are not changed.
The cacs_rate_audit attribute records whether the check was on
(enabled), the number of rate rows outside the ranges
(n_out_of_bounds), and those rows (rows). It is attached even when
the check is off (the default), with enabled = FALSE and no rows.
References
U.S. Census Bureau (2020). Understanding and Using American Community Survey Data: What All Data Users Need to Know. Chapter 8, Calculating Measures of Error for Derived Estimates.
See also
For the rates and their margins of error at more length, see
vignette("theory-derived-rates", package = "catchmentACS")
(https://joonho112.github.io/catchmentACS/articles/theory-derived-rates.html).
Other steps of the calculation:
cacs_acs_prefetch(),
cacs_intersect_weight(),
cacs_isochrone(),
cacs_propagate_moe(),
cacs_run()
Examples
# Turn the cache off while this example runs (see ?cacs_set_cache).
old <- options(catchmentACS.cache_enabled = FALSE)
# Example data bundled with the package: the drive-time areas are circles
# with a radius of 1 km per minute, and the ACS data are made up.
library(sf)
iso <- readRDS(system.file("extdata", "legacy_2025_isochrones.rds",
package = "catchmentACS"))
acs <- readRDS(system.file("extdata", "sample_alabama_subset.rds",
package = "catchmentACS"))
iso_07 <- iso[iso$site_id == "AL_SITE_07" & iso$drive_time_min == 10, ]
agg <- cacs_intersect_weight(iso_sf = iso_07, acs_sf = acs,
verbose = FALSE)
prop <- cacs_propagate_moe(agg, verbose = FALSE)
# The five rates, with margins of error at the 90 percent level
out <- cacs_derive_rates(prop, verbose = FALSE)
is_rate <- out$estimand_family == "derived_rate"
out[is_rate, c("variable", "estimate", "moe", "moe_formula_effective")]
#> # A tibble: 5 × 4
#> variable estimate moe moe_formula_effective
#> <chr> <dbl> <dbl> <chr>
#> 1 poverty_rate 0.602 0.0984 general_ratio_conservative
#> 2 snap_rate 0.0426 0.00557 general_ratio_conservative
#> 3 ssi_rate 0.165 0.0360 general_ratio_conservative
#> 4 unemp_rate 0.0630 0.0149 general_ratio_conservative
#> 5 labor_force_participation 0.846 0.112 general_ratio_conservative
# Margins of error with the default and with formula_dispatch = "auto"
out_auto <- cacs_derive_rates(prop, formula_dispatch = "auto",
verbose = FALSE)
tibble::tibble(variable = out$variable[is_rate],
moe_default = out$moe[is_rate],
moe_auto = out_auto$moe[is_rate],
formula_auto = out_auto$moe_formula_effective[is_rate])
#> # A tibble: 5 × 4
#> variable moe_default moe_auto formula_auto
#> <chr> <dbl> <dbl> <chr>
#> 1 poverty_rate 0.0984 0.0858 proportion_subset
#> 2 snap_rate 0.00557 0.00557 general_ratio_conservative
#> 3 ssi_rate 0.0360 0.0360 general_ratio_conservative
#> 4 unemp_rate 0.0149 0.0149 general_ratio_conservative
#> 5 labor_force_participation 0.112 0.0904 proportion_subset
options(old)