Skip to contents

catchmentACS 0.5.0 corrects the estimates of medians and per-person values and checks some arguments more strictly than 0.4 did. For code written for 0.3, vignette("porting-v03-to-v04", package = "catchmentACS") describes the changes made in 0.4.

library(catchmentACS)
library(dplyr)
library(sf)  # needed to subset the bundled sf objects with [

Medians and per-person values

cacs_run() averages the medians and per-person values of three tables of the American Community Survey (ACS) over the census tracts that a drive-time area overlaps: median household income (B19013_001), median home value (B25077_001), and per capita income (B19301_001). A median or per-person value from any other table is added up like a count. The weights of the average, called area shares, are proportional to the area of each tract’s overlap with the drive-time area and sum to one for each site, drive time, and variable. Before 0.5.0, they summed to one over all the variables of a site and drive time taken together. In a call with more than one variable, the estimates computed with these weights were therefore too small, and so were their margins of error, the half-widths of their confidence intervals (90 percent by default). When every tract has a row for each variable, they were smaller by a factor equal to the number of variables in the call. Counts and rates are computed with other weights and did not change.

Medians and per-person values computed with 0.4 need to be computed again. In a result, they are the rows whose weight_basis is "area_mean". The example below finds them in a run on made-up data installed with the package, in which the drive-time areas are circles and the ACS estimates are random numbers:

iso <- readRDS(system.file(
  "extdata", "legacy_2025_isochrones.rds", package = "catchmentACS"
))
acs <- readRDS(system.file(
  "extdata", "sample_alabama_subset.rds", package = "catchmentACS"
))
site_07 <- data.frame(site_id = "AL_SITE_07", lon = -85.365, lat = 31.655)

result <- cacs_run(
  site_07, state = "AL",
  precomputed_isochrones = iso[iso$site_id == "AL_SITE_07", ],
  acs = acs, verbose = FALSE
)
tibble::as_tibble(result) |>
  filter(weight_basis == "area_mean") |>
  select(drive_time_min, variable, estimate, moe)
#> # A tibble: 6 × 4
#>   drive_time_min variable   estimate    moe
#>            <int> <chr>         <dbl>  <dbl>
#> 1              5 B19013_001   68369  16626 
#> 2              5 B19301_001   16355   2422 
#> 3             10 B19013_001   68037. 16160.
#> 4             10 B19301_001   16725.  2356.
#> 5             15 B19013_001   60498.  7834.
#> 6             15 B19301_001   25135.  2643.

The year argument

cacs_acs_prefetch() accepts as year only a whole number from 2009 to 2024, and so does cacs_run() when it downloads the ACS estimates. In 0.4, they accepted any year from 2009 to the year before the current one and truncated a fractional year, such as 2023.5. In 0.5.0, a fraction and a year after 2024 give an error. When the estimates are supplied through acs, cacs_run() only records year.

Saved ACS estimates and the Census API key

cacs_acs_prefetch() reads estimates saved in the cache without a Census API key, as it did in 0.4. In 0.5.0, it looks for them before sending any request, where 0.4 first asked tidycensus for the list of ACS variables and went on if that failed. It needs the key only to download: when no saved result matches the call, or with force_refresh = TRUE.

A saved result matches only while the installed versions of R, catchmentACS, tidycensus, tigris, and sf stay the same (?cacs_acs_prefetch), so copies saved with 0.4 are not read. The first call with 0.5.0 downloads the estimates again and needs the key. Results that cacs_intersect_weight() saved in the cache with 0.4 are not reused either; they are computed again. From version 0.6.0 on, saved results are kept only until the R session ends unless a cache folder that lasts between sessions is set, and the folder used by versions 0.5.1 and earlier is not read (see ?cacs_cache_dir).

Sites as a plain data frame

cacs_isochrone() now also accepts the sites as a base R data frame with the columns site_id, lon, and lat, like site_07 above. In 0.4, it required an sf object of points or a tibble, and both are still accepted.

The isochrone column when cacs_run() builds the areas

With output = "list_column" or output = "both", the list-column form of cacs_run() has a column isochrone, which holds the drive-time area of each row as a one-row sf object. In 0.4, it held the areas only when they were supplied through precomputed_isochrones, and NULL when cacs_run() built them. In 0.5.0 it holds them in both cases. A message about the column is shown on every such call, even with verbose = FALSE.

Choices that are not implemented

weight_method = "population" is not implemented yet; area weighting ("area", the default) is the only method. In 0.5.0, cacs_run() gives the error before any step runs, with the class catchmentACS_error_credential. In 0.4, the error came at the intersection step, after the areas and the ACS estimates were ready, when bg_pop_sf was supplied; without bg_pop_sf, it came at the start, with the class catchmentACS_error_schema.

Choosing provider = "mapbox" or provider = "r5r" to build the areas, or a rates list other than cacs_acs_default_rates, the list of the five built-in rates, gives an error, as it did in 0.4.