18  A gated protocol for one real test

This is a protocol for producing auditable evidence, not an automatic estimator selector. The simulation guidance is exploratory and grid-specific, and this release carries 19 validated, in-support placements of its own. A lookup for a new test requires that test to earn the same two coordinates: its own standardized shape calibration (Phase C) and a reliability inside the grid’s support (Phase D).

flowchart TD
  A["Phase A · Name the reporting goal"] --> B["Phase B · Fit candidate item models and compute continuous rho-bar"]
  B --> C["Phase C · Run standardized observed-and-null shape calibration"]
  C --> D{"Shape validated and rho within nearest-tier support?"}
  D -- "no" --> E["Report descriptive consequences and the placement limit; no case verdict"]
  D -- "yes" --> F["Phase D · Exploratory sim-v3 lookup with tier gap and scope"]
  F --> G["Phase E · Fit goal-relevant candidates and a pre-specified comparator"]
  G --> H["Phase F · Independent-seed, dense-band, and cut-mass audits"]
  H --> I["Report case evidence and simulation evidence in separate registers"]

The gated protocol. A case-specific lookup is allowed only after standardized shape calibration and an in-support reliability placement.

18.1 Phase A — Name the target

Specify whether the report serves individual squared error, the finite-sample reported EDF, ranking/selection, or fixed cuts. A single estimate set need not optimize all targets. This choice is substantive and should be recorded before selecting a winner or comparator.

18.2 Phase B — Fit item models and measure continuous reliability

Fit each defensible item model and report the continuous \bar\rho used by the simulation. Here \bar\rho is an inverse-information approximation to MSEM, conditional on fitted item parameters. It substitutes an average of 1/J(\theta) and omits item-calibration and model uncertainty; it is not a universal reliability coefficient. The relation \operatorname{SD}(\widehat\theta_{PM})\approx\sqrt{\bar\rho} is a heuristic under constant-error normal–normal conditions, not an identity.

Keep marginal_rxx as a separate fitted-score functional. Compare gaps under the two coefficients without treating the coefficients themselves as commensurate. If model choice materially changes \bar\rho, shape evidence, or the case outputs, report that heterogeneity without assigning a causal mechanism to estimated discriminations.

18.3 Phase C — Rebuild the shape screen correctly

For both the observed empirical-histogram fit and every null refit:

  1. normalize the quadrature weights;
  2. center and scale the weighted support;
  3. resample on that standardized support;
  4. add 0.05 SD-unit jitter under a recorded seed; and
  5. compute the dip statistic.

Apply the full classifier, including mild-*, unclassified, and normal detectability branches. “No statistic exceeds its null” does not by itself mean normal. Freeze the code, seeds, observed statistics, null summaries, and quadrature artifacts before reading downstream case results.

If this calibration is unavailable, stop the case-specific lookup. The 47/54 warehouse flag count is evidence of model dependence, not a known false-positive rate.

18.4 Phase D — Locate only supported cells

Report N, continuous \bar\rho, nearest tier, absolute tier gap, and the nearest-tier support boundary. Do not silently clamp values outside support. The present simulation uses N=50,100,200,500, reliability tiers .5–.9, and a joint test-length plus c^* ladder. Its scope is dichotomous, unidimensional data with three studied shapes and one bank template/family.

Any lookup is exploratory and conditional on those scope assumptions. The seven MSEL safety labels must be reported when relevant, together with the fact that their H4 interpretation was formalized after outcomes in Addendum 006. They are not a blanket stop rule for every goal.

18.5 Phase E — Fit candidates and name the comparator

Use PM when individual squared error is the target. For distributional reporting, compare PM with CB and GR while stating their exact targets:

  • CB is the marginal-variance/posterior-independence specialization; it omits cross-person posterior covariance.
  • GR targets the posterior expected finite-sample EDF G_N and expected ranks, not an abstract population CDF.

Report both a pre-specified comparator and any selected-winner ratio. A ratio to the selected minimum answers a different question and is usually larger. Label the ratio convention—condition-mean or paired-geometric—because they are different estimands.

18.6 Phase F — Audit reproducibility and decision geometry

Refit independently and recompute the report. In this portfolio, PM and CB reproduce closely, while GR can reassign a nearly fixed quantile ensemble across people in dense near-rank bands on short Rasch forms. Exact ties are rare and are not the mechanism. Seed extraction as a safeguard, but do not mistake seeding for an independent-refit audit.

For selection, compute exact top-k membership with a declared boundary rule and report both method churn and independent-seed churn. For fixed cuts, report the share within a declared neighborhood of each cut and the 0.02-gap band census. These are diagnostics of local fragility, not proofs that everyone inside a band will cross.

18.7 Reporting contract

Keep two registers:

  1. Case evidence: observed consequence, seed sensitivity, and placement diagnostics.
  2. Simulation evidence: truth-based loss ratios, interval/label, estimand, grid scope, and post-outcome interpretation where applicable.

Do not convert the second into case advice when the first lacks a validated, in-support placement. This release demonstrates the reporting contract and the failure gates; it does not claim a fully validated automatic selector.