7  Nine combinations, four output families

This chapter specifies the comparison: what was fitted, what was recorded, and what counts as a difference worth reporting. Everything in it was fixed before the first production model ran.

7.1 The item models and the priors

Each case-cell is fitted under its item model, Rasch (\mathrm{logit}\,P(y_{pi}=1) = \theta_p - b_i) or two-parameter logistic (\lambda_i(\theta_p - b_i)), with the latent ability distribution G given one of three priors. The Gaussian arm is the convention and the comparator. The two flexible arms place a Dirichlet process mixture on G; they differ only in elicitation, the data-independent rule that sets the DP’s concentration behavior from the analysis sample size. The focused arm uses the calibrated elicitation the simulation carried as its flagship (mixture complexity concentrated on a handful of clusters); the broad arm is deliberately more diffuse, a sensitivity contrast. The theory volume (Lee, 2026) develops the construction; for this book it is enough that the DP arms can represent shoulders, heavy tails and multiple modes that the Gaussian cannot, at roughly fifteen times the computing cost (Chapter 8).

7.2 The summaries

Every fitted model is summarized three ways from the same posterior draws. That pairing reduces Monte Carlo variance in within-arm contrasts; it does not make the realized contrast noise-free. PM and CB are deterministic for fixed draws, while GR’s rank extraction can consume random numbers only when exact posterior-rank ties occur. The posterior mean (PM) minimizes each person’s expected squared error and shrinks toward the center; the ensemble of PM estimates is narrower than the population, by construction, in proportion to unreliability. The constrained Bayes summary (Ghosh, 1992) (CB) applies an affine marginal-variance rescaling. It matches the joint posterior moment target under posterior independence, but omits cross-person posterior covariances induced by shared items or hyperparameters; it repairs spread and cannot repair shape. The triple-goal summary (Shen & Louis, 1998) (GR) targets the posterior expected finite-sample realized EDF G_N and every person’s posterior expected rank, then reports the corresponding EDF quantile; it aims the ensemble at the distribution and the ordering jointly, with lower individual accuracy by design. Figure 11.1 in Chapter 11 shows all nine estimate sets on one real dataset; a reader who wants the summaries’ anatomy before the results may look ahead to it now.

Table 7.1: The three posterior summaries. Each is computed from the same fitted model; the optimality target is the quantity the summary is built to serve.
summary_id label source definition optimality_target
PM Posterior mean standard posterior expectation of theta_p individual squared-error loss
CB Constrained Bayes Ghosh (1992) posterior mean rescaled so the ensemble’s first two moments match the posterior expectation of the empirical moments ensemble distribution
GR Triple-goal Shen & Louis (1998) each estimate placed at the midpoint of its quantile interval of the estimated EDF ranks and distribution jointly

7.3 The four output families

The comparisons are organized by what a technical report prints, four families of output, because Part IV’s organizing result is that the same method contrast can be immaterial in one family and material in another on the same dataset.

F1, individual scores: per-person differences between two estimate sets, standardized by the default report’s SD, summarized by the median absolute movement (with p90 and maximum kept). F2, the reported distribution: the Kolmogorov–Smirnov distance between the two reported empirical distribution functions, plus the ratio of SDs, of IQRs, and the moments, kept separately because they can disagree (Chapter 11 shows they do, and why a single spread statistic misleads). F3, rankings and selection: rank correlations, percentile-rank movements, and the overlap of top groups (5, 10, 25%), reported as the share of the selected group that changes membership. F4, tails and cuts: the share of the sample beyond fixed cuts at \theta = 1.0, 1.5 and 2.0, as a signed change in percentage points and as a reclassification traffic count, plus percentile-defined decile membership with chance-corrected agreement.

Four contrasts cross the families. K1 compares every alternative combination with the default (Gaussian + PM), the practitioner’s actual question. K2 isolates the prior (DP-focused against Gaussian at fixed summary), K3 the summary (within the Gaussian arm unless stated), K4 the elicitation (focused against broad). K3 is paired on common draws and has much less Monte Carlo variation than K2 and K4, but it is not exact; every contrast is read against the paired independent-seed comparison.

7.4 Materiality

Each family has a pre-declared threshold above which a difference is material, worth acting on (Table 7.2). The thresholds encode report-scale judgments (a tenth of a standard deviation moves a percentile band; a five-point KS moves a norms table visibly; two percentage points at a cut moves a pass rate a program would notice), and they were fixed before fitting so that no threshold could be tuned to a result. The score and distribution thresholds aligned more closely with the observed contrasts than the tails threshold, which was aimed at net movement, missed the phenomenon that actually matters there (clump-edge jumps, Chapter 13), and the book says so where it reports it.

Table 7.2: Pre-declared materiality thresholds, per output family. Frozen before any model was fitted.
Family Material when
F1 Individual scores median |move| > 0.10 reference SD
F2 Reported distribution KS between reported EDFs > 0.05, or SD ratio outside [0.95, 1.05]
F3 Rankings and selection > 5% of ranks move > 2 percentile points, or top-decile Jaccard < 0.90
F4 Tails and cuts any fixed cut’s share moves > 2 percentage points

Scales are never standardized before comparison. The arms are identified on a common reporting scale (audited, not assumed, by a scale check on the PM ensembles), and standardizing would erase exactly the spread-and-shape differences F2 exists to measure; the cost of that refusal is that F1 movements are expressed in units of the default ensemble’s SD, which the definitions appendix (Appendix B) states precisely.

Ghosh, M. (1992). Constrained Bayes estimation with applications. Journal of the American Statistical Association, 87(418), 533–540. https://doi.org/10.1080/01621459.1992.10475236
Lee, J. (2026). Bayesian semiparametric item response modelling for person-specific latent traits: Theory, identification, and the literature behind the estimators. https://joonho112.github.io/dpmirt-theory-book/
Shen, W., & Louis, T. A. (1998). Triple-goal estimates in two-stage hierarchical models. Journal of the Royal Statistical Society Series B: Statistical Methodology, 60(2), 455–471. https://doi.org/10.1111/1467-9868.00135