A Simulation Study of Dirichlet Process Mixture Priors and Goal-Specific Posterior Summaries in Bayesian IRT

Design, results, and practical guidance from a preregistered study of 7,200 model fits

Author
Affiliation

JoonHo Lee

The University of Alabama

Published

August 9, 2026

About this book

Bayesian item response models are asked three kinds of question, about individuals, about rankings, and about population distributions, and operational practice often answers all three with one default: a Gaussian prior on ability, summarized by the posterior mean. This book reports a preregistered simulation study of two possible changes, flexible Dirichlet process mixture priors and goal-specific posterior summaries, crossed with each other and with the study’s information ladder, which moderates both changes in the simulated grid.

The study fits 7,200 models over 120 conditions: two model families (Rasch, 2PL), three latent shapes (normal, skewed, bimodal), five reliability tiers (0.5 to 0.9, achieved by nested test lengths plus cell-specific discrimination calibration), four sample sizes (50 to 500), three priors, and three summaries, with 20 replicate-paired datasets per condition and a locked analysis plan. Items are dichotomous and the latent model is unidimensional, with one item-bank template per model family; the reliability tiers therefore represent a joint length-plus-\(c^\star\) design rather than an isolated reliability intervention. All five primary hypotheses and both secondary hypotheses were supported, the negative control was quiet, and the preregistered safety screen produced the one result the study is obliged to carry alongside its headline: the flexible pipeline’s individual-accuracy penalty in bimodal low-reliability cells.

The findings, compressed. The flexible prior improves distribution recovery under non-normality by 20 percent at the design’s center and by half in the information-rich corner; normal controls mostly cluster near parity, with three resolved higher-loss cells. The posterior summary is the larger lever: swapping it moves distributional loss several times more than swapping the prior, and a Gaussian prior with the right summary beats a flexible prior with the wrong one in 100 of 120 cells. Position on the joint ladder helps organize the observed recipe, indexing what shape information reaches the prior (sharply for bimodality, mildly for skew) and setting the shrinkage the summaries repair. Among the studied estimators and conditions, rank loss is dominated by position on that ladder and varies little across estimator choices. The individual-accuracy tradeoff is also visible: seven bimodal cells at reliability 0.5 and 0.6 have 5 to 14 percent higher individual MSEL (ratios 1.052 to 1.138, every interval excluding 1); the complete cell-level table and Addendum 006’s post-outcome timing are reported in The safety result, stated as registered.

This is one volume of a three-part project, alongside a theory volume and a case-study volume; it is the volume that carries the evidence. It is written to be read at three depths: 7  Results at a glance alone for the answer, Part III for the results with their mechanics, and the whole book for the design, the machinery and the record. Named headline and audit quantities are computed at build time from frozen row-sets through the facts and table layers; documented protocol constants and archival quotations remain source-cited literals. Every figure ships with its generating script, and the preregistration record, including the two entries with disclosure duties, is reproduced in 15  The preregistered record.