Case Studies of Dirichlet Process Mixture Priors and Goal-Specific Posterior Summaries in Bayesian IRT
Thirteen real tests from the Item Response Warehouse: consequences, reproducibility, and what the companion simulation recommends
About this book
Every Bayesian item-response score report combines two choices: a prior for the latent ability distribution and a rule for reducing each person’s posterior to a reported number. Operational practice commonly uses a Gaussian prior and the posterior mean. The companion simulation studies how alternative priors and summaries recover known truth within a bounded grid; this volume asks what those choices change on real tests, where truth is unavailable.
The case study applies the same nine prior-summary combinations to thirteen assessments from the Item Response Warehouse, under Rasch and 2PL item models. It reports consequences for individual scores, reported distributions, rankings and selections, and fixed-cut shares. It also measures whether each reported summary reproduces under an independent seed. These are questions the real data can answer directly.
Transfer from the simulation is a separate question, and it is gated on a measurement this edition had to rebuild. An external audit of the previous edition found that the latent-shape screen computed its dip statistic on each fitted model’s native scale with a fixed absolute jitter, which made the observed statistic, its own null and the simulation’s reference condition mutually incommensurable. The supports needed to correct it had not been retained, so this edition refit: 26 observed densities and 5,158 null replicates, with the statistic standardized before jittering, applied identically to the cases, the nulls and the reference. The repair changes 1 label of 26, and that one change is reproduced by the old estimator as well, so it is refit variation at a borderline fit rather than a consequence of the correction. The shape coordinate is now validated on a commensurable standard, and 19 of the 26 case-cells carry both a determinate class and a reliability inside the simulation grid’s support. C12–Rasch and C4–2PL fall below that support and are excluded rather than clamped. The reliability side of the match remains an approximation, and every joined row is reported with its tier gap.
Within those limits, the descriptive results are still useful. Prior and summary swaps affect different output families; the ordering depends on the estimand and cannot be reduced to a universal larger-lever claim. Fixed-cut shares are usually stable, with the two largest exceptions localized at a dense band near the cut (C4–Rasch, 19.0 points; C2–Rasch, 10.6). Exact top-k membership exposes seed sensitivity that rank correlations conceal: Gaussian + PM changes as much as 28.0% of one cell’s selected decile across seeds. GR exceeds 0.10 common-reference SD movement on 11 of 24 short-form fits, including 10 with maximum \hat R<1.05; PM and CB remain within 0.028 SD throughout.
This is one volume of a three-part project. The theory volume defines the estimands and their assumptions; the simulation volume carries truth-based evidence within its studied grid; this case-study volume carries real-data consequence, measurement, reproducibility, and transfer-limit evidence. 19 What this release settles—and what it does not gives the qualified simulation findings, Part IV gives the case consequences, and 18 A gated protocol for one real test turns the validity conditions into a gated protocol. Headline facts are generated from this edition’s derived layer at render time. Input identities, post-outcome interpretations, unregistered analyses, and every correction are recorded in Appendix C — Computation and reproducibility record and Appendix D — Deviations, corrigenda, and the release record.