15 Each case, matched and judged
Part IV established consequence: the choices change real reports, by measured amounts, in family-specific patterns. This chapter adds the register the cases cannot supply. For each eligible case-cell, its coordinates select one cell of the simulation’s grid, and this chapter reads out what the simulation, where truth is known, concluded in that cell: which combination wins each inferential goal, by how much, and at what price to the goals it does not serve. Every sentence here is the simulation’s, transported; the cases contribute their location and nothing else.
The previous edition could not do this. Its shape coordinate rested on a statistic computed on an incommensurable scale (Section 5.4), so it carried the same join as an explicitly unvalidated sensitivity calculation and issued no case recommendations. The standardized refit has since closed that residual, and the labels it returns are, with one borderline exception, the labels the earlier editions used. What follows is therefore the same arithmetic the previous edition performed, now resting on a shape coordinate that has been validated rather than assumed.
15.1 The roster
Of 26 case-cells, 22 carry a determinate standardized shape class. Two of them, C12–Rasch (\bar\rho=.425) and C4–2PL (\bar\rho=.261), fall below the .5 tier’s nearest support boundary and are removed rather than clamped, leaving 19 eligible rows over 14 distinct simulation cells. Every row carries its continuous reliability, its nearest tier and the gap between them, which reaches 0.049 at worst and 0.033 at the median. That gap is the honest residual of this chapter: the shape coordinate is now measured on the simulation’s own standard, while the reliability coordinate is a fitted-model quantity matched to a design-time tier, and no refit can make those two constructions identical.
The simulation evidence is rebuilt directly from the current volume’s evidence.rds rather than copied from the first edition’s transfer table. On the primary distribution-recovery contrast, 14 of the 19 rows carry a strong_win label, with condition-mean loss ratios from 0.476 to 0.761. Five normal-consistent rows carry harm_flag labels close to parity (0.988 to 1.041), which is imprecision around nothing rather than harm: those are the calibration controls, seen from the case side.
15.2 Winners by goal
Figure 15.1 shows the six-combination condition-mean winner for four goals. A washed tile means the winner is within Monte Carlo error of its runner-up, so on the ranking column, where almost every tile is washed, the winner’s identity is descriptive rather than decisive. The case join adds a second layer of uncertainty on top of that, from the tier gap in the reliability coordinate.
A GR combination wins the distribution goal in 18 of 19 matched cells, and a PM combination wins individual-score MSEL in 19 of 19. The summary follows the goal almost without exception; the winning prior follows the case. This is evidence about the simulated grid, read at the cells these tests occupy, and it is the only register in this book that can say which choice is better rather than merely different.
15.3 Comparator-dependent loss ratios
Against the selected six-combination winner, Gaussian + PM has a distribution-loss ratio from 1.36 to 2.54, median 2.08. This is a post-selection ratio. Directly comparing Gaussian + PM with the pre-specified Gaussian + GR comparator gives a median ratio of 1.20 overall and 1.15 on the skewed and bimodal rows. The first calculation compares the default with a selected minimum; the second isolates the change from PM to GR under the Gaussian prior.
The prior and summary levers also depend on the estimand. In the 14 in-support non-normal rows, the summary-only swap beats the prior-only swap in only 3 rows. Their case-weighted geometric loss factors are 0.841 for summary-only and 0.811 for prior-only. By contrast, the simulation-wide full 3\times3 factorial decomposition assigns 82% to summary and 14% to prior. These are different estimands and are not presented as competing estimates of one quantity.
Figure 15.3 shows two full menus to make that dependence visible.
15.4 What this chapter establishes, and what it does not
For every eligible non-normal case-cell, the simulation’s verdict in the matching cell favors the flexible prior for distributional reporting, by enough to matter and with the GR summary doing the larger share of the work; for every matched cell of any shape, PM remains the summary for individual scores; and the default combination, judged where these cases live, roughly doubles distributional loss relative to the attainable.
Four limits travel with those statements. The join is unpreregistered and was assembled after both source volumes’ results, so it is exploratory in status even though its inputs are frozen. The reliability match is an approximation carrying a tier gap of up to 0.049. The ordering of the two levers depends on the estimand, as the comparator analysis above shows, so no universal larger-lever claim is made here. And nothing in this chapter is validated by the cases: a case study without truth cannot confirm a simulation, and the two-register discipline of Chapter 3 is what licenses the chapter at all. Its scope is dichotomous, unidimensional data, N from 50 to 500, the three simulated shapes, a joint test-length and c^* reliability ladder, and the studied item-bank template and estimator family.


