| Arm | Prior on the ability distribution | Truncation M | Sampler budget |
|---|---|---|---|
| Gaussian | theta_i ~ N(mu, sigma^2), the default in operational IRT | - | 8k/40k iterations (Rasch/2PL) |
| DP focused | DP mixture of normals; Gamma prior on the concentration calibrated so the prior expects few occupied clusters | 50 (75 at N = 500) | 12k/36k |
| DP broad | DP mixture of normals; deliberately diffuse concentration prior (sensitivity arm) | 50-150 by N | 12k/36k |
| Sources: sampler budgets are read from F$design (generated from data/design/sampler-budgets.csv); truncation schedules and prior descriptions are protocol constants from the frozen alpha-truncation manifest and production configuration. | |||
5 Estimands, losses, and contrasts
This chapter fixes the quantities the results chapters use: the nine estimator combinations, the five loss families and what each one measures, the common scale that makes losses comparable across arms, the aggregation ladder from replicate to condition, and the six registered contrasts with their evidence-label rules.
5.1 Nine estimators
| Summary | Decision problem it solves | What it preserves |
|---|---|---|
| PM (posterior mean) | minimize expected squared error for each individual ability | individual accuracy |
| CB (constrained Bayes) | minimize squared error subject to matching the mean and variance of the estimate set to the posterior expectations | first two moments of the ensemble |
| GR (triple-goal) | match the empirical distribution and ranks of the estimate set (Shen and Louis, 1998) | EDF and ranks of the ensemble |
The crossing of Table 5.1 and Table 5.2 yields the nine combinations that populate the winner maps. The summary is computed from retained posterior draws, so its cost is post-processing only; the arm changes the fitted model and its budget.
5.2 Five losses
| Family | Definition | Inferential goal | Confirmatory role |
|---|---|---|---|
| MSEL | mean squared error of individual ability estimates | individual scores | H4 safety |
| MSELR | mean squared error of percentile ranks (ties averaged) | ordering and selection | descriptive |
| KS | sup-distance between estimate and truth EDFs | distribution recovery | primary (H1-H3, H5, H6) |
| Quantile | weighted squared quantile gaps at deciles .10-.90 | distribution recovery | sensitivity |
| Tail | weighted squared gaps in shares beyond cutoffs -2, -1.5, 1.5, 2 | tail identification | descriptive; replicate-level registered ratios are floored |
| Source: preregistered loss definitions and roles, with replicate-level floor scope from Addendum 002 and H7 share-only floor scope from Addendum 003 amendment 001. | |||
Every loss compares an estimate set with the same replication’s realized true abilities, not with the population’s analytic form; the estimand is recovery of this sample’s ensemble, which is what an operational report is about. Figure 5.1 draws the three distributional losses on one low-reliability exemplar, where their differences are visible. The KS distance is the primary distributional loss by preregistration; the quantile loss squares and weights gaps at the five registered probabilities 0.10, 0.25, 0.50, 0.75 and 0.90, so it amplifies the same discrepancies. The tail loss compares lower-tail shares at negative cutoffs and upper-tail shares at positive cutoffs and is the one family whose comparator can be exactly zero, the origin of the floor discussed in Section 13.3. Squared rank-error uses percentile-scaled ranks with ties averaged. The individual loss (MSEL) is squared error in the reporting metric.
5.3 One scale for all arms
Losses are computed on a common reporting scale, the standardized scale of the generating population. In this simulation, the known generating item difficulties and scaled discriminations stored with each frozen dataset define the map from the constrained-item fit scale back to that DGP scale before losses are computed. This truth-based alignment is available because this is a simulation; it is not an identification procedure available in an operational analysis with unknown generating parameters. Without the explicit map, arm comparisons would mix scale conventions with substance. The mapping was repaired and re-verified after an external review found the earlier convention leaking across a boundary (Addendum 005; the incident and its verification are part of Appendix D), and every loss row carries the scale identifier and map hash it was computed under.
5.4 From replicate to condition
Two cell-level ratio estimands appear in the book, and the nonlinear aggregation means they are not interchangeable. For the registered evidence maps, losses are averaged first. If \(L_{A,cfr}\) is the loss for pipeline A in condition \(c\), form \(f\), and Monte Carlo replicate \(r\), define the equal-form condition mean
\[ \bar L_{A,c}=\frac{1}{F}\sum_f\left(\frac{1}{R_f}\sum_r L_{A,cfr}\right), \qquad \mathrm{RR}^{\mathrm{mean}}_c=\frac{\bar L_{A,c}}{\bar L_{B,c}}. \]
The frozen evidence column rr_point is \(\mathrm{RR}^{\mathrm{mean}}_c\). Its hierarchical bootstrap resamples forms and then paired replicates within form (\(B = 10,000\)). Registered cell labels and safety boundaries are attached to this ratio of equal-form condition means.
Some exploratory displays instead use the paired geometric-style cell ratio in ratio_cells:
\[ \mathrm{RR}^{\mathrm{paired}}_c =\exp\!\left[\frac{1}{F}\sum_f\left\{\frac{1}{R_f}\sum_r \log\left(\frac{L_{A,cfr}}{L_{B,cfr}}\right)\right\}\right]. \]
This second quantity preserves dataset pairing before aggregation. In general, \(\mathrm{RR}^{\mathrm{mean}}_c\ne\mathrm{RR}^{\mathrm{paired}}_c\); equal form weighting does not remove the distinction. Each ratio-using figure identifies which quantity it displays. Confirmatory hypotheses operate on the registered replicate-level log-ratio rows and pool them in mixed models with a form-family random intercept (Section 6.4); their coefficients are therefore not cell-level ratios of condition means either.
5.5 Six contrasts and the label rule
Six contrasts were registered: the primary distribution-recovery contrast (DP focused + GR against Gaussian + GR on KS), its broad-arm companion, the within-DP comparison, the MSEL safety contrast on the same combination pair, the CB-against-PM summary contrast, and the normal-cell control. Each cell of each contrast receives a locked evidence label: win if the bootstrap interval lies below 1 (strong win if below 0.95), tie if the interval sits inside [0.95, 1.05], and flag otherwise. The flag category is a residual, not a verdict; Section 8.6 decomposes it. The safety contrast additionally carries the 1.05 and 1.10 boundaries whose block label is the claim gate’s trigger (Section 6.5). The registered point ratios and evidence labels use rr_point, the ratio of condition means defined above. Exploratory displays that use ratio_cells name the paired geometric-style estimand explicitly; neither is described as a substitute for the other.
