19 What this release settles—and what it does not
This release’s strongest result is a reporting result: prior and summary choices can change different output families by materially different amounts on the same real dataset. A report that asks only whether person scores moved cannot certify its distribution, ranking, or cut decisions.
19.1 Consequence is output-specific
The frozen within-case contrasts remain informative without knowing truth. They show score movement, reported-EDF movement, selection churn, cut traffic, and independent-seed variation under declared metric definitions. Class-conditioned aggregates are grouped by the standardized shape class, which is measured against each case’s own null; a class records a detectable departure of a stated kind, not a claim that the population is exactly one of the three simulated shapes.
The structured normal-control profile is E-02 and remains exploratory. The cut/band mechanism is E-01 and remains exploratory. Item-model differences are fitted-model heterogeneity; the observed co-movements do not establish that discrimination parameters absorbed, preserved, or caused a true latent shape.
19.2 Reproducibility is a property of the reported quantity
Using the common Gaussian-PM cell SD, PM and CB remain within about 0.03 reference SD across the two independent fits. GR exceeds 0.10 on 11 of 78 fits, all short Rasch forms, while 10 of those failures have maximum \hat R<1.05. Exact ties are too rare to explain the result. The sorted GR ensemble is stable; dense near-rank perturbations reassign its quantiles across nearby people. Therefore chain diagnostics cannot substitute for an independent reproduction of the final report.
Selection inherits the same local fragility. Near-one rank correlations can coexist with meaningful top-group churn. This release distinguishes exact top-k rules from quantile-threshold rules and reports independent-seed churn as one paired reference. It is neither a full seed-to-seed uncertainty distribution nor evidence that the substantive selection is meaningless.
19.3 What sim-v3 contributes
Sim-v3 provides truth-based, goal-specific losses over its studied grid. The case book reconstructs those values from current authority files and exposes 19 in-support matched rows. Conditional on that roster, Gaussian + PM has median distribution loss 2.08 times the selected winner. Directly against Gaussian + GR, the median default-to-comparator ratio is 1.20 overall and 1.15 on the skewed and bimodal rows. The contrast shows why comparator selection must be explicit.
This join was assembled after both source volumes’ results, so it is unpreregistered and exploratory in status even though its inputs are frozen. The recommendation-eligible row count is 19. The shape coordinate is now measured on the simulation’s own standard; the reliability side remains an approximation. The simulation tiers were built by joint test length and c^* calibration from .5 to .9, fitted case coordinates and known DGP coordinates are not exactly commensurate, and every matched row carries a tier gap whose median is 0.033 and whose maximum is 0.049.
The seven MSEL safety labels remain relevant simulation evidence. Their H4 interpretation was formalized after outcomes in Addendum 006 and does not create a blanket distribution-recovery stop rule. C12–Rasch and C4–2PL are below nearest-tier reliability support and receive no cell verdict.
19.4 Scope and limitations
The simulation scope is dichotomous, unidimensional data; N=50–500; normal, strong positive skew, and strong bimodality; a joint length plus c^* reliability ladder; and the studied bank template and estimator family. The case portfolio is purposive, not population-representative, and contains no polytomous data. Its public-warehouse sources overrepresent shareable, curated, largely Western and low-stakes datasets.
\bar\rho is an inverse-information approximation conditional on fitted item parameters, not a full measure of calibration and model uncertainty. CB omits joint posterior covariances; GR targets a posterior expected finite-sample EDF, not an abstract population distribution. Density overlays compare a latent DP density with an empirical distribution of shrunken PM estimates and are explicitly non-like-for-like.
19.5 The next release gate
The standardized observed-and-null refit, the retained quadrature artifacts, and the newly frozen placement manifest that the previous edition set as its gate were delivered here, and the declared policy for reliabilities outside nearest-tier support is to decline extrapolation. What remains is the reliability coordinate itself. Closing the tier gap would take a simulation designed around fitted coordinates, and extending guidance below .45 would take a grid that reaches there. Until either exists, the durable deliverables are the audited consequence metrics, the matched simulation evidence with its gap reported, the reproducibility diagnostics, and the two-register reporting protocol.