1 Introduction
Assessment data are collected to answer more than one kind of question. A tutoring program wants each student’s ability estimate; an intervention study wants the five students most in need; an accountability report wants the share of a population below a benchmark. These are different estimation problems with different loss functions, yet operational practice typically answers all of them with a single set of numbers: point estimates from an item response model, fitted with a normal prior on ability and summarized by the posterior mean. Figure 1.1 illustrates the three questions on one simulated dataset.
Two distinct things can go wrong with the default. The first is prior misspecification: real ability distributions are sometimes skewed or multimodal (floor effects, mixed populations, instructional sorting), and a normal prior pulls every posterior toward a shape the population does not have. The literature has worried about this failure for decades and has produced a flexible remedy, the Dirichlet process mixture prior, which lets the data decide the shape. The second failure is less discussed: even with a correctly specified model, the posterior mean is not aligned with the distributional questions studied here. Posterior means minimize squared error person by person, and the same shrinkage that earns them that property makes their ensemble too narrow, so cutoff shares, tail counts and distributional shapes computed from them are biased even when every individual estimate is as good as it can be (Louis, 1984; Shen & Louis, 1998). Goal-specific summaries exist for exactly this reason: constrained Bayes corrects the ensemble’s spread (Ghosh, 1992), and the triple-goal estimator of Shen & Louis (1998) targets ranks and the distribution directly.
The two remedies act on different components, one on the model and one on what is computed from it, and they have almost never been studied together in item response settings. That gap is not an accident of attention. Studying them jointly requires a design in which the prior, the summary, and the amount of information the data carry all vary at once. In this study the third axis moderates how much either change helps: flexible-prior recovery depends on shape information transmitted by the likelihood, and rank recovery depends on distinctions among people supported by the data. Test reliability, the proportion of observed-score variance that is signal in its conventional use, motivates the third axis. The quantity implemented here is narrower: an inverse-information MSEM design coefficient, \(\operatorname{Var}(\theta)/[\operatorname{Var}(\theta)+E\{1/J(\theta)\}]\), used to index the average precision of a form for a generating population. It is not assumed to equal exact posterior/EAP or observed-score reliability. The design spans values near 0.5, typical of a short low-information instrument, through 0.9, typical of a much more informative form.
This book reports a simulation study built around that three-way interaction. Two model families (Rasch and 2PL), three latent shapes (normal, skewed, bimodal), five reliability tiers achieved by a nested-length ladder followed by cell-specific global discrimination calibration, and four sample sizes define 120 conditions; within each, every dataset is fitted under a Gaussian prior and two Dirichlet process mixture arms, and every fit is summarized by the posterior mean, constrained Bayes, and the triple-goal estimator, then scored on five loss functions matched to the three inferential goals. The models are unidimensional, the items are dichotomous, and each family uses one item-bank template. The resulting method rankings therefore require new evidence before transport to other dimensions, response formats, item banks, latent shapes, or information regimes. The comparisons of central interest were preregistered, with a locked analysis plan, a negative control, a safety test, and an amendment record whose timing is documented in Chapter 15. The study is the third build of its design; the two earlier versions died of defects that are documented in Appendix D, and the present version exists because their lessons were mechanical enough to enforce.
The reliability factor is consequently a realized length-plus-scale information ladder, not a randomized intervention on either length or discrimination alone. The confirmatory H6 contrast remains a valid test of moderation along that registered ladder; causal partitions between reliability, item count, global discrimination scale and the shape of the information curve are exploratory and not identified by this design.
Readers who want the conclusions before the machinery can hold four sentences from Chapter 7. Within this grid, the flexible prior improves distribution recovery most under non-normality with more information; normal controls mostly cluster near parity but include three resolved higher-loss cells. The posterior summary is usually the larger of the two studied levers. Position on the realized information ladder changes the recipe: which prior, which summary, and whether the choice matters materially. And the gains for populations accompany a measured, localized individual-accuracy penalty that this book reports with the same prominence as the gains, because its preregistration requires it to.
The plan of the book is as follows. Chapter 2 develops the two levers, shrinkage and its goal-specific repairs, and the Dirichlet process mixture prior, at the level of intuition needed to read the results; Chapter 3 does the same for reliability and its targeting, which is the design’s distinctive axis. Chapter 4 through Chapter 6 describe the design grid, the estimands and losses, the computational machinery, and the preregistered plan with its amendments. Chapter 7 through Chapter 15 present the results: the headline maps, the distribution-recovery findings, the lever comparison, the reliability recipes, the model-family dissociation, the individual-accuracy tradeoff, the goals that move little, evidence quality, and the complete preregistered record. Chapter 16 condenses practice into recipes, and Chapter 17 states scope and limitations. Appendices carry the computational record, reproducibility, condensed theory, the version history, and the full result tables.
