4  The design

The design crosses everything that the argument of Part I says must be crossed: model family, latent shape, reliability, sample size, prior, and summary. This chapter walks the grid, the replication structure inside each condition, and the two disclosed asymmetries that reliability-matching between families produces. Here reliability matching means the actual nested-length plus global-scale calibration, not length matching alone. Figure 4.1 shows the whole at a glance.

Design grid of 120 conditions with 20 replications each, alongside the multiplication chain from forms and draws to datasets, fits and loss values.
Figure 4.1: The design at a glance. Panel (a): the 120 conditions cross two model families, three latent shapes, five reliability tiers and four sample sizes; every tile holds 20 paired replications. Panel (b): within a condition, five independently drawn item banks (forms) each generate four response datasets; every dataset is fitted under all three priors, and each fit is summarized three ways and scored on five losses. The complete study is 2,400 datasets, 7,200 fits, and 115,200 loss values, with all comparisons replicate-paired within dataset.

4.1 Factors and their levels

Table 4.1: Design factors. The first four are crossed and define the 120 conditions; prior and summary vary within dataset and within fit. Source: tables/T-factors.rds.
Factor Levels Role
IRT model family Rasch; 2PL crossed, co-primary
Latent shape normal (control); skewed (k = 2.0); bimodal (delta = 0.90) crossed
Target reliability 0.5, 0.6, 0.7, 0.8, 0.9 (nested length + cell-specific c_star calibration) crossed
Sample size N 50, 100, 200, 500 crossed
Fitted prior Gaussian; DP focused; DP broad within dataset
Posterior summary PM; CB; GR (from the same fit) within fit
Form replicate 5 independent item banks per condition replication
Response draw 4 per form (20 datasets per condition) replication
Source: frozen production design as summarized in the generated F$design component; the reliability ladder uses the design's nested-length, cell-specific c_star calibration.

The table’s compact reliability label denotes the full frozen construction. Item count rises through a nested ladder, and a multiplier \(c^\star\) is then calibrated and applied separately for each model, form, tier and latent shape. There are 150 such calibration cells. Test length is constant across shapes within a model-by-form-by-tier subset, whereas \(c^\star\) is allowed to differ so the achieved inverse-information coefficient is matched.

Four decisions in Table 4.1 deserve their reasons. The three latent shapes are stylized rather than exhaustive: a standard normal as negative control, a strongly positively skewed population, and a strongly separated two-component mixture, each standardized to mean zero and unit variance so that shape, not location or scale, is what varies. The reliability tiers run low deliberately; 0.5 and 0.6 are unflattering territory for every method, and they are where short operational instruments live. Sample sizes stop at 500 because the earlier, non-confirmatory builds suggested that the qualitative map had stabilized by then, and the design prioritized replication (each condition’s N = 1,000 slot became five extra replicates). The two model families are co-primary rather than primary-and-robustness: every condition exists in both, and the family comparison of Chapter 11 is a designed contrast, not an afterthought.

4.2 Inside one condition

Replication has two layers because item banks are themselves random. Each condition draws five independent forms (item ladders calibrated to the tier as described in Section 3.3). A form supplies the same nested item subset and baseline difficulty geometry to the three shapes at a tier; the shape-specific \(c^\star\) multiplies its discriminations before responses are generated. Each form then generates four response datasets, giving 20 paired replications per condition; the between-form layer measures item-bank luck, and every condition-level number in this book weights the five forms equally. Each dataset is then fitted three times, once per prior arm, and each fit is summarized three ways and scored on five theta losses plus one item-parameter loss, so a condition contributes 60 fits and 960 loss values, and the study totals 2,400 datasets, 7,200 fits and 115,200 loss rows. All comparisons are replicate-paired: the same dataset underlies every arm and summary being compared, so between-dataset noise cancels from the contrasts.

4.3 Achieved reliability travels with the data

Table 4.2: The realized reliability ladder: analytic achieved reliability and mean test length per model family and tier. Source: tables/T-ladder.rds.
Target Achieved (analytic) Mean items Mean c_star Max abs. error
2PL
0.5 0.4999 7.4 0.9662 0.00200
0.6 0.6001 10.8 0.8854 0.00200
0.7 0.7000 16.4 0.8874 0.00094
0.8 0.7999 28.0 0.9191 0.00095
0.9 0.9001 62.2 0.9271 0.00110
Rasch
0.5 0.4999 8.0 0.9459 0.00087
0.6 0.6001 12.0 0.9101 0.00270
0.7 0.6997 18.0 0.9240 0.00180
0.8 0.8002 31.0 0.9161 0.00190
0.9 0.9000 68.2 0.9182 0.00064

A condition’s name carries its family, sample size, tier and shape; its data carry the measured reliability actually achieved, at form level and dataset level. The confirmatory models of Chapter 6 use the measured value, centered, rather than the tier label, so that the moderation analyses stand on realized information. Table 4.2 shows the item-count component of parity between families: at every tier the two families match achieved reliability to within 0.000 while differing in length by about 10 percent (Rasch needs the longer forms, lacking the discrimination parameters that concentrate information). In addition, the frozen \(c^\star\) values span 0.794–1.090 across the full calibration grid. Family comparisons are matched on the achieved scalar coefficient, not on equal item counts or identical discrimination scales. This is the first disclosed asymmetry of the family comparison; the second, differing sampler budgets, is described with the machinery (Section 6.2), and both are discussed where the families are compared (Section 11.3).

4.4 What the design does not vary

Reading the results requires knowing what was held fixed. Item-bank recipes (difficulty layouts, and baseline discrimination distributions for the 2PL) follow one balanced template per family; the population is always standardized; missingness, guessing, multidimensionality and person misfit are absent by construction; and priors on item parameters are common across arms, so the arms differ only in the ability prior. Each of these constants marks a boundary of scope, collected in Chapter 17, and one of them, the item-bank template, is the reason the form-replicate layer exists: within the template, banks still vary, and that variation is measured rather than assumed away.

What is not fixed is equally important. The global discrimination multiplier changes by cell, and matching \(\bar\rho_J\) does not make the full information curves identical across latent shapes. Tail-region information can differ even when the scalar average is matched. The design therefore supports paired prior-arm contrasts within cells and registered moderation along the realized ladder; it does not identify a pure causal effect of length, discrimination, or reliability considered separately.