Skip to contents

Scripts written for catchmentACS 0.1 or 0.2 can stop with an error under version 0.6.0, or run and give different results. vignette("porting-v03-to-v04", package = "catchmentACS") and vignette("porting-v04-to-v05", package = "catchmentACS") describe the changes of versions 0.4 and 0.5 in more detail.

Several errors about drive-time areas (isochrones), such as a missing column or the wrong coordinate reference system, refer to this article. The section on saved drive-time areas explains them.

library(catchmentACS)
library(sf)  # needed to subset the bundled sf objects with [

Two files that ship with the package stand in for data of your own. fixture_032_3site_fresh.rds holds the inputs and results of a run with version 0.2 for three points at public places in Birmingham, Mobile, and Huntsville, Alabama: 10-minute areas from the public OSRM (Open Source Routing Machine) demo server and American Community Survey (ACS) estimates for 2019–2023. The second file, legacy_2025_isochrones.rds, holds circles around 20 made-up sites, used here as drive-time areas that pass the current checks.

# Inputs and results of a run with version 0.2
old_run <- readRDS(system.file(
  "testdata", "fixture_032_3site_fresh.rds",
  package = "catchmentACS", mustWork = TRUE
))

# Drive-time areas saved with version 0.2
old_areas <- old_run$iso_sf_v2

# Drive-time areas that pass the current checks
example_areas <- readRDS(system.file(
  "extdata", "legacy_2025_isochrones.rds",
  package = "catchmentACS", mustWork = TRUE
))

Summary of changes

Version Change Effect on an older script What to do
0.3 Drive-time areas need a ring_topology column equal to "cumulative" Saved areas give an error Add the column, or join bands with cacs_rings_to_cumulative()
0.3 cacs_isochrone() joins the bands that the osrm package returns for several drive times into cumulative areas Saved areas built with OSRM for several drive times may be bands, and the estimates for their longer drive times then cover only a band Check them as below; build them again, or join them with cacs_rings_to_cumulative() after adding isomin and isomax
0.3, 0.4 Without res, OSRM areas use 30L with osrm_mode = "demo" (the default) and 70L with osrm_mode = "docker" (50 in version 0.2) New areas differ from the old ones Set res = 50L for the grid of version 0.2 (iso_args = list(res = 50L) in cacs_run())
0.3 ssi_rate was B19056_001 / B11001_001 and is now B19056_002 / B19056_001 Old values cannot be compared, and ACS data downloaded with the default variables of version 0.2 give NA Download the ACS data again
0.5 The weights that average medians and per-person values sum to one for each variable Old values from a call with more than one variable are too small: when every tract has all of them, an old value is the correct one divided by the number of variables Compute them again
0.4 The rows of rates_breakdown in summary() are no longer shifted Old means, standard deviations, and counts of missing values belong to other rates Compute them again
0.4 tibble::as_tibble(), and functions that call it, put the rows of the five rates first within each site and drive time In versions 0.4 to 0.5.1, rows picked by position after tibble::as_tibble() differed, and dplyr::semi_join(), dplyr::anti_join(), and summaries after dplyr::rowwise() could give wrong results without a warning Nothing from version 0.6.0 on, which keeps the order again. rate_first = TRUE, or options(catchmentACS.rate_first_default = TRUE) for calls from your own code, puts the rate rows first
0.3 A result saved in the cache needs a checksum file Results saved by versions 0.1 and 0.2 are not used Nothing; the old cache folder can be deleted by hand (see ?cacs_cache_dir)
0.2 cacs_acs_prefetch() leaves out tracts numbered 9900 or higher (water and special-purpose tracts) and tracts with no area Fewer tracts, with a message drop_water_tracts = FALSE keeps them in a new download; data saved in the cache without them need force_refresh = TRUE
0.2, 0.3 Results gain the columns n_tracts_num, n_tracts_den, and ring_topology Columns move Pick columns by name

Drive-time areas saved with versions 0.1 and 0.2

cacs_intersect_weight() checks the drive-time areas it is given, and so does cacs_run() with precomputed_isochrones, after it downloads the ACS estimates if acs is not supplied. The areas must be an sf object in EPSG:4326 (longitude and latitude) with the 16 columns of a cacs_isochrone() result, including ring_topology, which versions 0.1 and 0.2 did not have. cacs_validate_iso() makes the same checks of the columns, their types, and the coordinate reference system, and returns one row for each problem, with a line of R code for a fix in its example column. Checking saved areas with it first avoids a download that ends in an error:

issues <- cacs_validate_iso(old_areas)
issues[, c("check", "col", "actual")]
#> # A tibble: 1 × 3
#>   check              col           actual 
#>   <chr>              <chr>         <chr>  
#> 1 iso_missing_column ring_topology missing
writeLines(issues$example)
#> iso_sf$ring_topology <- "cumulative"

This fix is right only if each row holds the whole area reachable within its drive_time_min, so that each area contains the shorter ones of the same site. cacs_validate_iso() flags rows whose isomin is above 0, but it does not compare the areas. These areas have one drive time for each site, so there is nothing to compare.

Drive-time bands, also called rings, cover only the minutes between two drive times, so a band shares no area with the band inside it. When several drive times are requested in one call, the osrm package returns bands, and since version 0.3 cacs_isochrone() joins them into cumulative areas. Saved areas that versions 0.1 and 0.2 built with OSRM for several drive times may be bands. The share of a shorter area that lies inside the longer one tells bands from cumulative areas, as in this example of one site with two bands:

make_square <- function(xmin, ymin, xmax, ymax) {
  sf::st_polygon(list(rbind(
    c(xmin, ymin), c(xmax, ymin), c(xmax, ymax),
    c(xmin, ymax), c(xmin, ymin)
  )))
}
core <- make_square(-87.10, 33.00, -87.00, 33.10)

# Two bands: 0 to 5 minutes, and 5 to 10 minutes (a larger square
# with the first one cut out)
bands <- sf::st_sf(
  site_id = "S01",
  isomin = c(0L, 5L),
  isomax = c(5L, 10L),
  geometry = sf::st_sfc(
    core,
    sf::st_difference(make_square(-87.20, 32.90, -86.90, 33.20), core),
    crs = 4326
  )
)

# Share of the shorter area that lies inside the longer one
inside_share <- function(shorter, longer) {
  overlap <- sf::st_intersection(sf::st_geometry(longer),
                                 sf::st_geometry(shorter))
  sum(as.numeric(sf::st_area(overlap))) /
    sum(as.numeric(sf::st_area(shorter)))
}
inside_share(shorter = bands[1, ], longer = bands[2, ])
#> [1] 0

cacs_rings_to_cumulative() joins each band with the bands inside it, using the columns isomin and isomax (the start and end of each band, in minutes):

cumulative <- cacs_rings_to_cumulative(bands)
sf::st_drop_geometry(cumulative)
#>   site_id isomin isomax drive_time_min ring_topology
#> 1     S01      0      5              5    cumulative
#> 2     S01      0     10             10    cumulative
inside_share(shorter = cumulative[1, ], longer = cumulative[2, ])
#> [1] 1

The result still lacks the other columns of a cacs_isochrone() result. For areas made by another tool, these columns can hold the values that cacs_isochrone() gives an area built without a problem, with the meanings given in the Value section of ?cacs_isochrone. provider and provider_requested must each be one of the four routing service names, "osrm", "ors", "mapbox", or "r5r", and provider is copied into the result, so for such areas it is only a label:

areas <- cumulative
areas$provider <- "osrm"
areas$provider_requested <- "osrm"
areas$provider_downgrade <- FALSE
areas$profile <- "car"
areas$routing_engine_version <- "osrm-pkg/unknown"
areas$polygon_simplification_tolerance <- NA_real_
areas$osm_snapshot_date <- "unknown"
areas$osm_snapshot_status <- "unknown_best_effort"
areas$generated_at <- Sys.time()
areas$isochrone_empty <- FALSE
areas$failure_reason <- NA_character_
areas$retry_count <- 1L

nrow(cacs_validate_iso(areas))
#> [1] 0

For example, cacs_run() uses an area only when its isochrone_empty is FALSE and its failure_reason is NA. An isochrone_empty of NA or a failure_reason of "" passes the checks, but the area then counts as a routing failure.

cacs_intersect_weight() takes areas only in EPSG:4326, and cacs_validate_iso() reports areas in another system:

bad_crs <- sf::st_transform(example_areas[1:2, ], 4269)
cacs_validate_iso(bad_crs) |>
  dplyr::filter(check == "iso_crs") |>
  dplyr::select(check, actual, expected, example)
#> # A tibble: 1 × 4
#>   check   actual    expected  example                                 
#>   <chr>   <chr>     <chr>     <chr>                                   
#> 1 iso_crs EPSG:4269 EPSG:4326 iso_sf <- sf::st_transform(iso_sf, 4326)

isochrone_empty and provider_downgrade must be logical, and drive_time_min and retry_count integers:

bad_type <- example_areas[1, ]
bad_type$provider_downgrade <- NA_character_
issues <- cacs_validate_iso(bad_type)
issues[, c("col", "actual", "expected")]
#> # A tibble: 1 × 3
#>   col                actual      expected          
#>   <chr>              <chr>       <chr>             
#> 1 provider_downgrade <character> logical TRUE/FALSE
writeLines(issues$example)
#> iso_sf$provider_downgrade <- FALSE

Running an older analysis again

The run saved in fixture_032_3site_fresh.rds shows what changes when an analysis from version 0.2 is run again on its saved inputs. With the areas and the ACS estimates supplied, cacs_run() builds no areas and downloads nothing:

old_sites <- sf::st_drop_geometry(old_run$sites_sf)
old_sites <- old_sites[, c("site_id", "lon", "lat")]
old_run_areas <- old_run$iso_sf_v2
old_run_areas$ring_topology <- "cumulative"

rerun <- cacs_run(
  old_sites,
  state = "AL",
  drive_times = 10,
  precomputed_isochrones = old_run_areas,
  acs = old_run$acs_sf,
  weight_args = list(keep_tract_audit = TRUE),
  verbose = FALSE
)
#> Warning: Some rates are `NA` or use a replacement formula:
#> • `failure_origin = "carrier"`: 3 rate rows, `NA` because a count or margin of
#>   error that the rate needs is missing.
#> ℹ The failure_origin and moe_fallback_reason columns give the reason on each
#>   rate row.

The warning gives one line for each reason found among the rate rows, and here there is one: the rows whose failure_origin is "carrier", rates that are NA because the estimate for the numerator or the denominator, or its margin of error (the half-width of its 90 percent confidence interval), is missing. Here they are the ssi_rate rows, one for each site. The new results compare with the saved ones as follows:

compare <- dplyr::inner_join(
  tibble::as_tibble(old_run$result_run_v2) |>
    dplyr::select(site_id, variable, saved = estimate),
  tibble::as_tibble(rerun) |>
    dplyr::select(site_id, variable, new = estimate),
  by = c("site_id", "variable")
)
dplyr::filter(
  compare,
  site_id == "AL_BHM_01",
  variable %in% c("B01003_001", "B19013_001", "B19301_001",
                  "poverty_rate", "ssi_rate")
)
#> # A tibble: 5 × 4
#>   site_id   variable         saved       new
#>   <chr>     <chr>            <dbl>     <dbl>
#> 1 AL_BHM_01 B01003_001   64233.    64233.   
#> 2 AL_BHM_01 B19013_001    5301.    68912.   
#> 3 AL_BHM_01 B19301_001    3949.    51332.   
#> 4 AL_BHM_01 poverty_rate     0.192     0.192
#> 5 AL_BHM_01 ssi_rate         1        NA
# All other counts and rates at the three sites
same <- dplyr::filter(
  compare,
  !variable %in% c("B19013_001", "B19301_001", "ssi_rate")
)
all.equal(same$saved, same$new)
#> [1] TRUE

The counts and four of the rates, 45 rows in all, did not change. Median household income (B19013_001) and per capita income (B19301_001) grew by the same factor wherever they have a value:

proxies <- dplyr::filter(compare, variable %in% c("B19013_001", "B19301_001"))
dplyr::mutate(proxies, ratio = new / saved)
#> # A tibble: 6 × 5
#>   site_id   variable   saved    new ratio
#>   <chr>     <chr>      <dbl>  <dbl> <dbl>
#> 1 AL_BHM_01 B19013_001 5301. 68912.  13.0
#> 2 AL_BHM_01 B19301_001 3949. 51332.  13.0
#> 3 AL_HSV_01 B19013_001 5674. 73757.  13  
#> 4 AL_HSV_01 B19301_001 3332. 43322.  13.0
#> 5 AL_MOB_01 B19013_001   NA     NA   NA  
#> 6 AL_MOB_01 B19301_001 1921. 24975.  13.0

The factor, 13, is the number of ACS variables in the run, in which every tract has a row for each variable. Before version 0.5.0, the weights that average medians and per-person values summed to one over all the variables of a site and drive time; vignette("porting-v04-to-v05", package = "catchmentACS") describes the correction.

In the saved results, ssi_rate is 1 at every site, because version 0.2 divided B19056_001 by B11001_001, two counts of all households. It now divides the households with Supplemental Security Income in the past 12 months, B19056_002, by all households, B19056_001. The saved ACS estimates lack B19056_002:

cacs_acs_default_rates$ssi_rate
#>          num          den 
#> "B19056_002" "B19056_001"
setdiff(unlist(cacs_acs_default_rates), old_run$acs_sf$variable)
#> [1] "B19056_002"

A new download with the default variables of cacs_acs_prefetch() or cacs_run() includes it, since it is in cacs_acs_default_vars.

Tables made with summary() in version 0.2 are also affected. Until version 0.4, each row of its rates_breakdown element showed the mean, standard deviation, and count of missing values of the next rate in alphabetical order, and the last row those of the first rate. Version 0.4 corrected this and added per-site tables of the rates, which cacs_summary_as_markdown() formats as Markdown.

Since version 0.3, keep_tract_audit = TRUE, passed above to cacs_intersect_weight() through weight_args, adds the attribute cacs_tract_audit. It has one row for each tract in each drive-time area, with the tract’s coverage weight, the share of its area inside the drive-time area. It can replace code that intersects the areas with tract boundaries to list the tracts:

tracts <- attr(rerun, "cacs_tract_audit")
names(tracts)
#> [1] "site_id"        "drive_time_min" "GEOID"          "area_wt"       
#> [5] "int_area_m2"    "tract_area_m2"
dplyr::count(tracts, site_id, drive_time_min, name = "tracts")
#> # A tibble: 3 × 3
#>   site_id   drive_time_min tracts
#>   <chr>              <int>  <int>
#> 1 AL_BHM_01             10     37
#> 2 AL_HSV_01             10     42
#> 3 AL_MOB_01             10     49

Building drive-time areas again

Areas built again can differ from the saved ones. One reason is the OSRM grid resolution res: the osrm package draws each area from travel times to a grid of res by res points around the site. Without res, version 0.2 used 50; version 0.6.0 uses 30L with osrm_mode = "demo", the default, and 70L with osrm_mode = "docker". Version 0.2 dropped res from iso_args with a warning, so a script that gave it there used 50 as well. To build the areas on the grid of version 0.2, set res = 50L in cacs_isochrone(), or iso_args = list(res = 50L) in cacs_run(). ?cacs_isochrone describes the grid, an option that changes the values, and the waiting time on the demo server. The road data of the server and the coordinates of the sites can change too.

Saved results in the cache

Since version 0.3, a result saved in the cache is used only when the checksum file next to it (.fingerprint) is present and its checksum matches. Results saved by versions 0.1 and 0.2 have no checksum file, so the first run after the upgrade builds the areas and downloads the ACS estimates again. Versions 0.5.1 and earlier kept saved results in the user cache folder of the operating system, which version 0.6.0 neither reads nor deletes; ?cacs_cache_dir gives the path of that folder on each system, and the folder can be deleted by hand.

cacs_set_cache(FALSE) turns the cache off for the rest of the R session, so that a script uses no saved results. The option catchmentACS.cache_dir or the environment variable CACS_CACHE_DIR sets another cache folder (see cacs_cache_dir()). Results saved in the default folder, inside the temporary folder of the R session, are deleted when the session ends.