2  What the simulation established

This chapter compresses the companion simulation volume (Lee, 2026) to the findings this book leans on. Readers of that volume can skip to Chapter 3; readers who want the design, the preregistration registry, or the full evidence maps should treat this chapter as a pointer to it. The simulation fitted 7,200 models over 120 conditions: two item models (Rasch and 2PL), three latent shapes (normal, strongly skewed, strongly bimodal), five reliability tiers (\bar\rho from .5 to .9, achieved by calibrated test length together with a cell-specific discrimination multiplier c^*), four sample sizes (50 to 500), and the three priors, each fit summarized three ways. All five primary hypotheses and both secondary hypotheses were supported under a locked analysis plan.

2.1 The flexible prior works where it should, and only there

Under non-normality, replacing the Gaussian prior with the focused DPM improves recovery of the ability distribution by about 20% at the design’s center (a Kolmogorov–Smirnov loss ratio of .80), and the advantage deepens by a factor of .89 with every doubling of the sample. Under normality it ties: in the forty normal control cells the flexible prior won none, at a measured flexibility cost of about 1.6% of distribution-recovery accuracy. The two DP elicitations, focused and broad, differ by about one percent, an amount the simulation’s own sensitivity analysis labels specification-dependent; nothing in this book turns on the distinction, and the case results will show the same (Chapter 12).

2.2 The lever answer depends on the estimand

The simulation-wide full 3\times3 factorial decomposition attributes 82% of within-cell distribution-loss variation to summary and 14% to prior. That is a factorial sums-of-squares estimand, not the same quantity as one summary-only swap versus one prior-only swap. Across the in-support non-normal matched rows, the latter comparison does not support a universal “summary larger” claim (Chapter 15).

The proposed mechanism is shrinkage: PM compresses the reported ensemble, while GR is designed to restore a distributional target. In the crossed simulation comparison, Gaussian + GR has lower paired-geometric loss than DP-focused + PM in 100 of 120 cells. Starting from Gaussian + PM in non-normal cells, the paired-geometric single-swap factors are 0.746 for summary-only, 0.879 for prior-only, and 0.593 for both. These ratios are reported with their convention; they are not interchangeable with factorial shares.

Within the studied grid the goal split is strong: PM is favored for individual squared error and GR is generally favored for distribution recovery. This is bounded grid evidence, not an equivalence theorem or a claim about all estimators.

2.3 Reliability conditions everything

Both levers answer to the reliability of the form, for different reasons. The summaries repair shrinkage, and the amount of shrinkage is set by reliability (a PM ensemble’s SD is about \sqrt{\bar\rho} times the population’s), so the summary swap matters most on unreliable forms. The prior can only model what the test transmits. Lower reliability reduces the visible bimodal signal, while larger N can increase exploitable shape information; the simulation does not license the absolute claim that no number of examinees can restore modes. The registered reliability-moderation slope decomposes accordingly, into -0.775 for the bimodal cells against -0.253 for the skewed ones: the reliability gradient is mostly a bimodality gradient. For rank accuracy, two separate marginal regressions gave R^2\approx.94 for reliability tier and R^2<0.0001 for estimator combination. They are not additive variance components or an equivalence test; they describe small estimator differences within this grid.

The simulation volume also gives an exploratory, grid-specific guidance table. It is reproduced for context, not applied as case advice in this release.

Table 2.1: Exploratory, grid-specific sim-v3 guidance. It is not a preregistered general protocol and is not applied as case advice in this release.
Goal Exploratory grid guidance Status
Individual scores Posterior mean; any prior (Gaussian suffices) Exploratory; dichotomous, unidimensional, N 50-500, three shapes, joint length+c* ladder
Ranking / selection Estimator differences were small in this grid; reliability was the dominant marginal predictor Exploratory; dichotomous, unidimensional, N 50-500, three shapes, joint length+c* ladder
Distribution, reliability up to ~.7 Gaussian + GR (add the DP prior only for skew at large N) Exploratory; dichotomous, unidimensional, N 50-500, three shapes, joint length+c* ladder
Distribution, reliability .8-.9, non-normal plausible DP (focused) + GR; check individual-score loss if scores are also reported Exploratory; dichotomous, unidimensional, N 50-500, three shapes, joint length+c* ladder
Defensibly normal, or mixed goals with no shape diagnosis Gaussian + goal-matched summary Exploratory; dichotomous, unidimensional, N 50-500, three shapes, joint length+c* ladder

2.4 The safety corner

In seven cells, all bimodal at reliability .5 and .6 with N of 200 or 500, the flexible pipeline charges individuals 5 to 14% in squared error (every interval excluding parity), and the charge grows with N: more data sharpen the prior’s conviction about the modes without sharpening any single person’s likelihood, so people between the modes are pulled toward the nearer one with increasing confidence. The interpretation and H4 scope restriction were recorded after outcomes in Addendum 006. That addendum does not create a blanket withholding rule for distribution recovery. C12–Rasch was formerly mapped to this corner, but its \bar\rho=.425 lies below the simulation’s nearest support; the match was removed in v3 and stays removed here (Chapter 17).

2.5 Rasch and 2PL

The simulation ran every condition under both item models as co-primary family. Its conclusions transfer in structure between them (the same winner geometry, the same goal split, quiet normal columns in both), with one registered asymmetry: the flexible prior’s advantage is deeper for bimodal populations under Rasch and deeper for skewed populations under the 2PL. The proposed mechanism, labeled an interpretation rather than an established result, is that the Rasch model’s sum-score sufficiency transmits a direct image of the ability distribution, which preserves a bimodal trace, while the 2PL’s estimated discriminations blur that image but absorb part of a skewed population’s score asymmetry into the likelihood. On real data the item model turns out to matter one step earlier than this, at the point of measuring what shape a dataset has at all; that is Chapter 5, and the two threads meet in Chapter 16.

Lee, J. (2026). A simulation study of Dirichlet process mixture priors and goal-specific posterior summaries in Bayesian IRT. https://joonho112.github.io/dpmirt-simulation-study/