5 Three coordinates, measured twice
To use the simulation’s answer, a real test must first be located on the simulation’s grid, and the grid has three axes: sample size, reliability, and latent shape. Sample size sounds like the easy one and nearly is; the other two are measurements, made under a model, with error and with model-dependence that turn out to be findings in their own right. This chapter is about the measuring. Its summary, drawn in Figure 5.1, is that a real test does not occupy one exact cell of the grid. Every case has two fitted coordinates, one per item model, and the shape class differs between the two models for six of the thirteen cases. The shape coordinate reported here is the one this edition rebuilt on a standardized statistic (Section 5.4); the reliability coordinate is a fitted-model quantity matched to the grid’s nearest tier, and it carries a gap that no refit removes.
5.1 Sample size, and the one hard constraint
The design uses two samples per case. Coordinates are measured on a characterization sample of up to 5,000 respondents, which never touches the method comparison; the nine combinations are then fitted to an analysis sample of at most 500, drawn once with a recorded seed. The cap is a package constraint, not a choice: the DPM elicitation routine accepts at most 500 persons, and the elicitation is a function of the analysis sample size. Three small datasets are fitted whole; every large case contributes exactly 500 (C13: 497 after three all-null respondents were dropped at fit time, a detail recorded because the design frame says 500). For grid placement, N matches downward to the nearest simulation value, so a 341-person case reads the N = 200 cell.
5.2 Reliability has no single value
Reliability enters the project twice, as a screening variable from the warehouse and as the placement coordinate, and the first edition of this book planned to treat the two as interchangeable. They are not, and the discrepancy is one of the sharpest methodological findings the portfolio produced. The warehouse’s coefficient is EAP-based (mirt’s marginal_rxx); the simulation’s grid is built on \bar\rho, an inverse-information approximation to mean squared error of measurement, conditional on fitted item parameters. It averages 1/J(\theta) and omits item-calibration and model uncertainty. Both are useful fitted-model functionals, but they do not agree about the portfolio: the median 2PL-minus-Rasch gap is +0.215 under marginal_rxx against +0.004 under \bar\rho, the two coefficients agree in sign for only 6 of 13 cases, and on the vocabulary-checklist case C4 they move in opposite directions by half a scale point each (+0.48 against -0.51) on the same pair of fits. Figure 5.2 displays the disagreement.
The lesson is not that one coefficient is wrong. The EAP-based coefficient answers a question about score precision under the fitted scoring rule; the inverse-information approximation answers a question about information against the latent metric, and a 2PL that concentrates its discrimination on a few items can raise the first while lowering the second. The lesson is operational: “the test’s reliability” is underdetermined until the functional is named, and a reader placing their own test on the simulation’s grid must compute the grid’s own functional, \bar\rho, under the model they intend to fit. The warehouse coefficient screened this corpus; it places nothing.
5.3 Shape cannot be read off a small sample
Latent shape is the coordinate the whole flexible-prior question turns on, and it is the one that cannot be read safely from a small sample. A naive fixed-threshold screen flagged 47 of 54 small warehouse datasets under the 2PL and 22 under Rasch. Truth is unknown in those real datasets, so these are not measured false-positive rates. They instead show that the naive flag is strongly model-dependent and likely contaminated by finite-design artifacts. Figure 5.3 shows the relevant behavior under known normality.
The screen calibrates a null per case: two hundred refits of data simulated from a normal population with the case’s own n, item count and item parameters, giving each statistic its own 95th percentile under normality for exactly this design. A case earns a class only by exceeding that null and by reaching at least half the simulation condition’s departure, so the class means “departs, and by enough for the simulation’s cells to speak to it”. The classifier has two further branches that matter and that the first two editions of this book failed to describe. Between a quarter and a half of the condition it returns mild-skewed or mild-bimodal. And it asserts normal only when this design’s nulls are tight enough that a condition-sized departure of either kind would have been detected; otherwise it returns undetermined, because “not flagged” and “normal” are not synonyms.
5.4 The statistic that had to be rebuilt
An external audit of the previous edition found a load-bearing defect in one of the three statistics. The dip was computed by resampling from the fitted density and adding jitter of 0.05 on each fit’s native scale, and the fitted density’s standard deviation ranges from 0.51 to 2.15 across the 26 case-cells. The smoothing was therefore between 2.3% and 9.7% of a standard deviation depending on the case, and the case’s statistic, its own null (generated at the Gaussian fit’s scale) and the simulation’s reference condition (computed on standardized draws) were three mutually incompatible quantities. Skewness, excess kurtosis and the KS distance were never affected: they are computed analytically from the density weights as standardized moments, so they are scale-invariant by construction. The defect was confined to the dip.
Because the fitted supports and weights behind the original null draws had not been retained, the repair required refitting rather than recomputation. This edition ran it: 26 observed empirical-histogram refits and 5,158 retained null replicates of 5,200 attempted, reproducing the original fits exactly – same characterization samples, same convergence ladder, same seed tree – and changing one thing. The dip is now computed by the published ancestor’s procedure, which standardizes the weighted support to mean zero and unit variance and then jitters on that scale, and the same estimator is applied to the observed densities, to every null replicate and to the simulation’s reference conditions.
Two checks establish that the refit is faithful and that the repair is the only thing that moved. The scale-invariant statistics reproduce the frozen values to within 0.00018 in skewness on 24 of the 26 cells. The two that move further are numerically borderline density fits rather than disagreements about the data: C12–Rasch, which needs the 8,000-cycle rung of the convergence ladder, lands 0.064 away, and C9–Rasch, at the portfolio’s largest fitted spread, lands 0.005 away. The reference condition’s dip, in turn, is 0.054 unjittered against 0.054 under the standardized estimator, so the jitter asymmetry the audit flagged in the half-condition gate was immaterial in size.
The labels themselves barely moved. Exactly 1 of 26 changes: C9–Rasch, which the previous screen left undetermined, now reaches the KS branch and is reported as non-normal without a class. That single change is not attributable to the standardization. Running the old native-scale estimator through the same machinery reproduces 25 of the 26 frozen labels and moves C9–Rasch in the same direction, so what moved it is the refit variation of a borderline fit. All nine skewed, all seven bimodal and all five normal-consistent labels stand under both estimators.
The correction was necessary and its consequence for the portfolio is nearly nil, which is the most useful outcome a repair of this kind can have. The shape coordinate is now measured on a standard commensurable with its own null and with the simulation grid it is matched to, and everything the previous editions reported by shape class survives intact.
One inference the repair does not license is worth stating, because it was our own first expectation. The dip’s correlation with the fitted latent standard deviation falls only from 0.59 to 0.54. Standardization removes the incomparability between a statistic and its reference; it does not remove the association between bimodality and variance, and most of that association is signal rather than artifact, since a population with two separated modes genuinely has a larger variance.
Two safeguards in the screen are worth carrying into practice, and neither would have repaired a statistic computed on inconsistent scales: require both design-specific evidence and a departure of practically comparable magnitude, and read skew, dip and KS jointly rather than separately.
5.5 Reliability-tier placement and the freeze
The continuous fitted \bar\rho values and analysis seeds were frozen before the production fits. The first edition then discretized reliability to the nearest tier but silently mapped every value below .55 to .5. C12–Rasch (\bar\rho=.425) and C4–2PL (\bar\rho=.261) lie below the .5 tier’s nearest-cell support boundary of .45; their former .5 matches are removed, and this edition reports for every row the continuous coordinate, its nearest tier, the absolute gap and the support status.
That leaves 19 rows carrying both a validated shape class and a reliability inside the grid’s support. Those rows are matched to simulation cells in Chapter 15. The reliability side of the match remains an approximation and is the weaker of the two coordinates: the case’s \bar\rho is a fitted-model quantity computed on real responses, while the grid’s tier is a design-time quantity computed on known generating parameters, and the two are commensurable in definition rather than identical in construction. Every matched row is therefore reported with its tier gap, which reaches 0.049 at worst.



