Converts a formula specification and data frame into a validated list suitable for passing to Stan. Constructs design matrices with optional standardisation, builds group indices, normalises survey weights, and computes all required dimensions.
Usage
prepare_stan_data(
formula,
data,
weights = NULL,
stratum = NULL,
psu = NULL,
state_data = NULL,
prior = NULL,
center = TRUE,
scale = TRUE
)Arguments
- formula
An object of class
"hbb_formula"(fromhbb_formula()), or a raw formula that will be parsed viahbb_formula().- data
A data frame containing provider-level variables: the response, trials, fixed effects, and optionally the grouping variable.
- weights
Character string naming the column in
datacontaining survey sampling weights, orNULL(default) for unweighted analysis.- stratum
Character string naming the column in
datacontaining sampling stratum identifiers, orNULL.- psu
Character string naming the column in
datacontaining primary sampling unit (PSU) identifiers, orNULL.- state_data
A data frame of group-level (state-level) variables for policy moderators. Required if
formula$policyis notNULL. Must contain a column matching the grouping variable for merging.- prior
A prior specification list, or
NULL(default) to use package defaults. (Prior specification API is under development.)- center
Logical. If
TRUE(default), subtract column means from non-intercept columns of X and V. The means are stored in the output for back-transformation.- scale
Logical. If
TRUE(default), divide non-intercept columns of X and V by their standard deviations after centering. The SDs are stored in the output for back-transformation.
Value fields
NInteger. Number of observations.
PInteger. Number of provider covariates including intercept.
SInteger. Number of groups (states). 1 if no grouping.
QInteger. Number of policy covariates including intercept.
N_posInteger. Number of positive observations (z == 1).
KInteger. Random effects dimension per group.
yInteger vector of length N. Response counts.
n_trialInteger vector of length N. Trial sizes.
zInteger vector of length N. Participation indicators.
XNumeric matrix N x P. Design matrix (column 1 = intercept).
stateInteger vector of length N. Group indices 1..S.
VNumeric matrix S x Q, or
NULLif no groups.idx_posInteger vector. Row indices where z == 1.
w_tildeNumeric vector of normalised weights, or
NULL.stratum_idxInteger vector of stratum indices, or
NULL.psu_idxInteger vector of PSU indices, or
NULL.n_strataInteger. Number of unique strata, or
NULL.n_psuInteger. Number of unique PSUs, or
NULL.x_centerNamed numeric vector of column means, or
NULL.x_scaleNamed numeric vector of column SDs, or
NULL.v_centerNamed numeric vector of V column means, or
NULL.v_scaleNamed numeric vector of V column SDs, or
NULL.formulaThe
hbb_formulaobject.priorThe prior specification (or
NULL).group_levelsCharacter vector of original group labels.
model_typeCharacter. One of
"base","weighted","svc","svc_weighted".
See also
hbb_formula(), validate_hbb_data()
Other data:
validate_hbb_data()
Examples
# Minimal example with the small synthetic dataset
data(nsece_synth_small, package = "hurdlebb")
f <- hbb_formula(y | trials(n_trial) ~ poverty + urban)
d <- prepare_stan_data(f, nsece_synth_small)
d
#> Hurdle Beta-Binomial Data
#> -------------------------
#> Observations (N) : 504
#> Covariates (P) : 3 (intercept + 2 predictors)
#> Groups (S) : 1
#> Policy vars (Q) : 1
#> Positive obs : 313 (62.1%)
#> Zero rate : 0.379
#> RE dimension (K) : 0
#> Model type : base
#> Survey weights : no
#> X standardised : yes
# With grouping and weights
f2 <- hbb_formula(
y | trials(n_trial) ~ poverty + urban + (1 | state_id)
)
d2 <- prepare_stan_data(f2, nsece_synth_small, weights = "weight")
d2
#> Hurdle Beta-Binomial Data
#> -------------------------
#> Observations (N) : 504
#> Covariates (P) : 3 (intercept + 2 predictors)
#> Groups (S) : 51
#> Policy vars (Q) : 1
#> Positive obs : 313 (62.1%)
#> Zero rate : 0.379
#> RE dimension (K) : 2
#> Model type : weighted
#> Survey weights : yes (range 0.05--14.92)
#> X standardised : yes
# Full SVC model with policy moderators
data(nsece_state_policy, package = "hurdlebb")
f3 <- hbb_formula(
y | trials(n_trial) ~ poverty + urban +
(poverty + urban | state_id) +
state_level(mr_pctile)
)
d3 <- prepare_stan_data(
f3, nsece_synth_small,
weights = "weight", state_data = nsece_state_policy
)
d3
#> Hurdle Beta-Binomial Data
#> -------------------------
#> Observations (N) : 504
#> Covariates (P) : 3 (intercept + 2 predictors)
#> Groups (S) : 51
#> Policy vars (Q) : 2
#> Positive obs : 313 (62.1%)
#> Zero rate : 0.379
#> RE dimension (K) : 6
#> Model type : svc_weighted
#> Survey weights : yes (range 0.05--14.92)
#> X standardised : yes