Updating code written for versions 0.1 and 0.2
Source:vignettes/porting-v01-to-v03.Rmd
porting-v01-to-v03.RmdScripts written for catchmentACS 0.1 or 0.2 can stop with an error
under version 0.6.0, or run and give different results.
vignette("porting-v03-to-v04", package = "catchmentACS")
and
vignette("porting-v04-to-v05", package = "catchmentACS")
describe the changes of versions 0.4 and 0.5 in more detail.
Several errors about drive-time areas (isochrones), such as a missing column or the wrong coordinate reference system, refer to this article. The section on saved drive-time areas explains them.
library(catchmentACS)
library(sf) # needed to subset the bundled sf objects with [Two files that ship with the package stand in for data of your own.
fixture_032_3site_fresh.rds holds the inputs and results of
a run with version 0.2 for three points at public places in Birmingham,
Mobile, and Huntsville, Alabama: 10-minute areas from the public OSRM
(Open Source Routing Machine) demo server and American Community Survey
(ACS) estimates for 2019–2023. The second file,
legacy_2025_isochrones.rds, holds circles around 20 made-up
sites, used here as drive-time areas that pass the current checks.
# Inputs and results of a run with version 0.2
old_run <- readRDS(system.file(
"testdata", "fixture_032_3site_fresh.rds",
package = "catchmentACS", mustWork = TRUE
))
# Drive-time areas saved with version 0.2
old_areas <- old_run$iso_sf_v2
# Drive-time areas that pass the current checks
example_areas <- readRDS(system.file(
"extdata", "legacy_2025_isochrones.rds",
package = "catchmentACS", mustWork = TRUE
))Summary of changes
| Version | Change | Effect on an older script | What to do |
|---|---|---|---|
| 0.3 | Drive-time areas need a ring_topology
column equal to "cumulative"
|
Saved areas give an error | Add the column, or join bands with
cacs_rings_to_cumulative()
|
| 0.3 |
cacs_isochrone() joins the bands that the
osrm package returns for several drive times into cumulative areas |
Saved areas built with OSRM for several drive times may be bands, and the estimates for their longer drive times then cover only a band | Check them as below; build them again, or join them
with cacs_rings_to_cumulative() after adding
isomin and isomax
|
| 0.3, 0.4 | Without res, OSRM areas use
30L with osrm_mode = "demo" (the default) and
70L with osrm_mode = "docker" (50 in version
0.2) |
New areas differ from the old ones | Set res = 50L for the grid of version 0.2
(iso_args = list(res = 50L) in
cacs_run()) |
| 0.3 |
ssi_rate was
B19056_001 / B11001_001 and is now
B19056_002 / B19056_001
|
Old values cannot be compared, and ACS data downloaded
with the default variables of version 0.2 give NA
|
Download the ACS data again |
| 0.5 | The weights that average medians and per-person values sum to one for each variable | Old values from a call with more than one variable are too small: when every tract has all of them, an old value is the correct one divided by the number of variables | Compute them again |
| 0.4 | The rows of rates_breakdown in
summary() are no longer shifted |
Old means, standard deviations, and counts of missing values belong to other rates | Compute them again |
| 0.4 |
tibble::as_tibble(), and functions that
call it, put the rows of the five rates first within each site and drive
time |
In versions 0.4 to 0.5.1, rows picked by position after
tibble::as_tibble() differed, and
dplyr::semi_join(), dplyr::anti_join(), and
summaries after dplyr::rowwise() could give wrong results
without a warning |
Nothing from version 0.6.0 on, which keeps the order
again. rate_first = TRUE, or
options(catchmentACS.rate_first_default = TRUE) for calls
from your own code, puts the rate rows first |
| 0.3 | A result saved in the cache needs a checksum file | Results saved by versions 0.1 and 0.2 are not used | Nothing; the old cache folder can be deleted by hand
(see ?cacs_cache_dir) |
| 0.2 |
cacs_acs_prefetch() leaves out tracts
numbered 9900 or higher (water and special-purpose tracts) and tracts
with no area |
Fewer tracts, with a message |
drop_water_tracts = FALSE keeps them in a
new download; data saved in the cache without them need
force_refresh = TRUE
|
| 0.2, 0.3 | Results gain the columns n_tracts_num,
n_tracts_den, and ring_topology
|
Columns move | Pick columns by name |
Drive-time areas saved with versions 0.1 and 0.2
cacs_intersect_weight() checks the drive-time areas it
is given, and so does cacs_run() with
precomputed_isochrones, after it downloads the ACS
estimates if acs is not supplied. The areas must be an
sf object in EPSG:4326 (longitude and latitude) with the 16
columns of a cacs_isochrone() result, including
ring_topology, which versions 0.1 and 0.2 did not have.
cacs_validate_iso() makes the same checks of the columns,
their types, and the coordinate reference system, and returns one row
for each problem, with a line of R code for a fix in its
example column. Checking saved areas with it first avoids a
download that ends in an error:
issues <- cacs_validate_iso(old_areas)
issues[, c("check", "col", "actual")]#> # A tibble: 1 × 3
#> check col actual
#> <chr> <chr> <chr>
#> 1 iso_missing_column ring_topology missing
writeLines(issues$example)#> iso_sf$ring_topology <- "cumulative"
This fix is right only if each row holds the whole area reachable
within its drive_time_min, so that each area contains the
shorter ones of the same site. cacs_validate_iso() flags
rows whose isomin is above 0, but it does not compare the
areas. These areas have one drive time for each site, so there is
nothing to compare.
Drive-time bands, also called rings, cover only the minutes between
two drive times, so a band shares no area with the band inside it. When
several drive times are requested in one call, the osrm package returns
bands, and since version 0.3 cacs_isochrone() joins them
into cumulative areas. Saved areas that versions 0.1 and 0.2 built with
OSRM for several drive times may be bands. The share of a shorter area
that lies inside the longer one tells bands from cumulative areas, as in
this example of one site with two bands:
make_square <- function(xmin, ymin, xmax, ymax) {
sf::st_polygon(list(rbind(
c(xmin, ymin), c(xmax, ymin), c(xmax, ymax),
c(xmin, ymax), c(xmin, ymin)
)))
}
core <- make_square(-87.10, 33.00, -87.00, 33.10)
# Two bands: 0 to 5 minutes, and 5 to 10 minutes (a larger square
# with the first one cut out)
bands <- sf::st_sf(
site_id = "S01",
isomin = c(0L, 5L),
isomax = c(5L, 10L),
geometry = sf::st_sfc(
core,
sf::st_difference(make_square(-87.20, 32.90, -86.90, 33.20), core),
crs = 4326
)
)
# Share of the shorter area that lies inside the longer one
inside_share <- function(shorter, longer) {
overlap <- sf::st_intersection(sf::st_geometry(longer),
sf::st_geometry(shorter))
sum(as.numeric(sf::st_area(overlap))) /
sum(as.numeric(sf::st_area(shorter)))
}
inside_share(shorter = bands[1, ], longer = bands[2, ])#> [1] 0
cacs_rings_to_cumulative() joins each band with the
bands inside it, using the columns isomin and
isomax (the start and end of each band, in minutes):
cumulative <- cacs_rings_to_cumulative(bands)
sf::st_drop_geometry(cumulative)#> site_id isomin isomax drive_time_min ring_topology
#> 1 S01 0 5 5 cumulative
#> 2 S01 0 10 10 cumulative
inside_share(shorter = cumulative[1, ], longer = cumulative[2, ])#> [1] 1
The result still lacks the other columns of a
cacs_isochrone() result. For areas made by another tool,
these columns can hold the values that cacs_isochrone()
gives an area built without a problem, with the meanings given in the
Value section of ?cacs_isochrone. provider and
provider_requested must each be one of the four routing
service names, "osrm", "ors",
"mapbox", or "r5r", and provider
is copied into the result, so for such areas it is only a label:
areas <- cumulative
areas$provider <- "osrm"
areas$provider_requested <- "osrm"
areas$provider_downgrade <- FALSE
areas$profile <- "car"
areas$routing_engine_version <- "osrm-pkg/unknown"
areas$polygon_simplification_tolerance <- NA_real_
areas$osm_snapshot_date <- "unknown"
areas$osm_snapshot_status <- "unknown_best_effort"
areas$generated_at <- Sys.time()
areas$isochrone_empty <- FALSE
areas$failure_reason <- NA_character_
areas$retry_count <- 1L
nrow(cacs_validate_iso(areas))#> [1] 0
For example, cacs_run() uses an area only when its
isochrone_empty is FALSE and its
failure_reason is NA. An
isochrone_empty of NA or a
failure_reason of "" passes the checks, but
the area then counts as a routing failure.
cacs_intersect_weight() takes areas only in EPSG:4326,
and cacs_validate_iso() reports areas in another
system:
bad_crs <- sf::st_transform(example_areas[1:2, ], 4269)
cacs_validate_iso(bad_crs) |>
dplyr::filter(check == "iso_crs") |>
dplyr::select(check, actual, expected, example)#> # A tibble: 1 × 4
#> check actual expected example
#> <chr> <chr> <chr> <chr>
#> 1 iso_crs EPSG:4269 EPSG:4326 iso_sf <- sf::st_transform(iso_sf, 4326)
isochrone_empty and provider_downgrade must
be logical, and drive_time_min and retry_count
integers:
bad_type <- example_areas[1, ]
bad_type$provider_downgrade <- NA_character_
issues <- cacs_validate_iso(bad_type)
issues[, c("col", "actual", "expected")]#> # A tibble: 1 × 3
#> col actual expected
#> <chr> <chr> <chr>
#> 1 provider_downgrade <character> logical TRUE/FALSE
writeLines(issues$example)#> iso_sf$provider_downgrade <- FALSE
Running an older analysis again
The run saved in fixture_032_3site_fresh.rds shows what
changes when an analysis from version 0.2 is run again on its saved
inputs. With the areas and the ACS estimates supplied,
cacs_run() builds no areas and downloads nothing:
old_sites <- sf::st_drop_geometry(old_run$sites_sf)
old_sites <- old_sites[, c("site_id", "lon", "lat")]
old_run_areas <- old_run$iso_sf_v2
old_run_areas$ring_topology <- "cumulative"
rerun <- cacs_run(
old_sites,
state = "AL",
drive_times = 10,
precomputed_isochrones = old_run_areas,
acs = old_run$acs_sf,
weight_args = list(keep_tract_audit = TRUE),
verbose = FALSE
)#> Warning: Some rates are `NA` or use a replacement formula:
#> • `failure_origin = "carrier"`: 3 rate rows, `NA` because a count or margin of
#> error that the rate needs is missing.
#> ℹ The failure_origin and moe_fallback_reason columns give the reason on each
#> rate row.
The warning gives one line for each reason found among the rate rows,
and here there is one: the rows whose failure_origin is
"carrier", rates that are NA because the
estimate for the numerator or the denominator, or its margin of error
(the half-width of its 90 percent confidence interval), is missing. Here
they are the ssi_rate rows, one for each site. The new
results compare with the saved ones as follows:
compare <- dplyr::inner_join(
tibble::as_tibble(old_run$result_run_v2) |>
dplyr::select(site_id, variable, saved = estimate),
tibble::as_tibble(rerun) |>
dplyr::select(site_id, variable, new = estimate),
by = c("site_id", "variable")
)
dplyr::filter(
compare,
site_id == "AL_BHM_01",
variable %in% c("B01003_001", "B19013_001", "B19301_001",
"poverty_rate", "ssi_rate")
)#> # A tibble: 5 × 4
#> site_id variable saved new
#> <chr> <chr> <dbl> <dbl>
#> 1 AL_BHM_01 B01003_001 64233. 64233.
#> 2 AL_BHM_01 B19013_001 5301. 68912.
#> 3 AL_BHM_01 B19301_001 3949. 51332.
#> 4 AL_BHM_01 poverty_rate 0.192 0.192
#> 5 AL_BHM_01 ssi_rate 1 NA
# All other counts and rates at the three sites
same <- dplyr::filter(
compare,
!variable %in% c("B19013_001", "B19301_001", "ssi_rate")
)
all.equal(same$saved, same$new)#> [1] TRUE
The counts and four of the rates, 45 rows in all, did not change.
Median household income (B19013_001) and per capita income
(B19301_001) grew by the same factor wherever they have a
value:
proxies <- dplyr::filter(compare, variable %in% c("B19013_001", "B19301_001"))
dplyr::mutate(proxies, ratio = new / saved)#> # A tibble: 6 × 5
#> site_id variable saved new ratio
#> <chr> <chr> <dbl> <dbl> <dbl>
#> 1 AL_BHM_01 B19013_001 5301. 68912. 13.0
#> 2 AL_BHM_01 B19301_001 3949. 51332. 13.0
#> 3 AL_HSV_01 B19013_001 5674. 73757. 13
#> 4 AL_HSV_01 B19301_001 3332. 43322. 13.0
#> 5 AL_MOB_01 B19013_001 NA NA NA
#> 6 AL_MOB_01 B19301_001 1921. 24975. 13.0
The factor, 13, is the number of ACS variables in the run, in which
every tract has a row for each variable. Before version 0.5.0, the
weights that average medians and per-person values summed to one over
all the variables of a site and drive time;
vignette("porting-v04-to-v05", package = "catchmentACS")
describes the correction.
In the saved results, ssi_rate is 1 at every site,
because version 0.2 divided B19056_001 by
B11001_001, two counts of all households. It now divides
the households with Supplemental Security Income in the past 12 months,
B19056_002, by all households, B19056_001. The
saved ACS estimates lack B19056_002:
cacs_acs_default_rates$ssi_rate#> num den
#> "B19056_002" "B19056_001"
#> [1] "B19056_002"
A new download with the default variables of
cacs_acs_prefetch() or cacs_run() includes it,
since it is in cacs_acs_default_vars.
Tables made with summary() in version 0.2 are also
affected. Until version 0.4, each row of its
rates_breakdown element showed the mean, standard
deviation, and count of missing values of the next rate in alphabetical
order, and the last row those of the first rate. Version 0.4 corrected
this and added per-site tables of the rates, which
cacs_summary_as_markdown() formats as Markdown.
Since version 0.3, keep_tract_audit = TRUE, passed above
to cacs_intersect_weight() through
weight_args, adds the attribute
cacs_tract_audit. It has one row for each tract in each
drive-time area, with the tract’s coverage weight, the share of its area
inside the drive-time area. It can replace code that intersects the
areas with tract boundaries to list the tracts:
#> [1] "site_id" "drive_time_min" "GEOID" "area_wt"
#> [5] "int_area_m2" "tract_area_m2"
dplyr::count(tracts, site_id, drive_time_min, name = "tracts")#> # A tibble: 3 × 3
#> site_id drive_time_min tracts
#> <chr> <int> <int>
#> 1 AL_BHM_01 10 37
#> 2 AL_HSV_01 10 42
#> 3 AL_MOB_01 10 49
Building drive-time areas again
Areas built again can differ from the saved ones. One reason is the
OSRM grid resolution res: the osrm package draws each area
from travel times to a grid of res by res
points around the site. Without res, version 0.2 used 50;
version 0.6.0 uses 30L with
osrm_mode = "demo", the default, and 70L with
osrm_mode = "docker". Version 0.2 dropped res
from iso_args with a warning, so a script that gave it
there used 50 as well. To build the areas on the grid of version 0.2,
set res = 50L in cacs_isochrone(), or
iso_args = list(res = 50L) in cacs_run().
?cacs_isochrone describes the grid, an option that changes
the values, and the waiting time on the demo server. The road data of
the server and the coordinates of the sites can change too.
Saved results in the cache
Since version 0.3, a result saved in the cache is used only when the
checksum file next to it (.fingerprint) is present and its
checksum matches. Results saved by versions 0.1 and 0.2 have no checksum
file, so the first run after the upgrade builds the areas and downloads
the ACS estimates again. Versions 0.5.1 and earlier kept saved results
in the user cache folder of the operating system, which version 0.6.0
neither reads nor deletes; ?cacs_cache_dir gives the path
of that folder on each system, and the folder can be deleted by
hand.
cacs_set_cache(FALSE) turns the cache off for the rest
of the R session, so that a script uses no saved results. The option
catchmentACS.cache_dir or the environment variable
CACS_CACHE_DIR sets another cache folder (see
cacs_cache_dir()). Results saved in the default folder,
inside the temporary folder of the R session, are deleted when the
session ends.