| Record | Timing | Content |
|---|---|---|
| Preregistration (locked 2026-07-24) | before the production wave | estimands, H1-H5, one Holm family, inclusion policy, thresholds |
| Addendum 001 | 2026-07-26: after P-2; before P-3 and production | reference grid for H1: which reading of 'at mean logN' the contrast uses |
| Addendum 002 | 2026-07-27: after P-2; before P-3 and production | symmetric floor on both operands of registered replicate-level log ratios; non-positive losses are floored, not dropped; condition-level ratios are outside its scope |
| Addendum 003 (+ amendment 001) | 2026-07-29/30: after P-3; before any production fit | separate H6-H7 Holm pair; amendment: H7 uses a share-only floor and form_family clustering (form_rep sensitivity) |
| Addendum 004 | 2026-07-31: 1,798 fits complete; no analyzer or outcome | optional direction-blind replication extension; the locked wording ladder was not added here (trigger did not fire) |
| Addendum 005 | 2026-08-03: post-fit, before production analysis | common reporting scale repair plus restriction of H6 to the realized reliability/item-count/information ladder |
| Addendum 006 | 2026-08-06: after production outcomes were known | post-outcome reporting policy for H4 safety: full 7-cell anchor in the confirmatory record; concise warnings and cross-references permitted in later synthesis |
| Source: preregistration and Addenda 001-006 registry. Dates, the 1,798-fit state at Addendum 004, and amendment timing are historical registry facts, not analytical outcomes recomputed from data/tidy. | ||
15 The preregistered record
Earlier chapters reported registered results where their content belonged. This chapter collects the complete confirmatory record in the form specified by the preregistration, together with the registration history itself. It consolidates source-backed results and timing rather than introducing a new analysis. The record establishes which results retain confirmatory status; it does not require later descriptive or exploratory synthesis to imitate the confirmatory presentation.
15.1 What was locked, and when
The analysis plan was locked on 2026-07-24, before the production wave ran, and its SHA-256 fingerprint travels in every confirmatory output row. The locked content includes the replicate-level paired log loss-ratio estimand (2,400 rows, 20 per condition), the linear mixed model with its fixed structure and form-family random intercept, measured centered reliability as the moderator, one Holm family over exactly five p-values at \(\alpha = 0.05\), the non-inferiority margin log(1.05) for H4, the opportunity region (non-normal shapes at N of 200 and 500), the tie band [0.95, 1.05] and safety boundaries 1.05 and 1.10, the inclusion policy (every completed fit enters; diagnostics gate wording, not membership), and the refusal rule under which a result that cannot be computed as registered is reported as such rather than approximated. The registration history after lock consists of the records in Table 15.1. The timing column distinguishes three states that should not be collapsed: a pilot outcome may already exist, fitting may already be under way, and production analysis may still not exist.
Read chronologically, Addenda 001 and 002 follow P-2 but precede P-3 and production. Addendum 003 follows P-3, discloses that the pilot prompted H6 and H7, and precedes every production fit; its Amendment 001 specifies H7’s share-only floor and form_family clustering. Addendum 004 was logged with 1,798 production fits complete, but before a production analyzer or outcome existed. It added an optional direction-blind replication extension, not the wording ladder already present in the lock. Addendum 005 is post-fit and pre-analysis: it repairs the reporting scale and restricts H6 to the realized reliability/item-count/information ladder. Addendum 006 alone is post-outcome.
The Addendum 006 policy is applied proportionately in this book. The full seven-cell table, ratios and intervals appear with the confirmatory H4 record here and with the main H4 result in Section 12.1. Later summaries may state the bimodal, low-reliability exception concisely and cross-reference one of those anchors. They need not reproduce seven rows every time H4 is mentioned. That policy keeps the local harm visible without turning a post-outcome reporting rule into a constraint on unrelated distributional claims or free exploratory synthesis.
15.2 The primary family
| Hypothesis | Estimate | SE | 95% CI | p (Holm) | Decision |
|---|---|---|---|---|---|
| H1 flexible prior improves distribution recovery (non-normal) | -0.221 | 0.013 | [-0.246, -0.195] | 1.35e-28 | supported |
| H2 the advantage grows with sample size | -0.114 | 0.004 | [-0.123, -0.106] | 1.24e-131 | supported |
| H3 model-by-shape dissociation | -0.127 | 0.023 | [-0.175, -0.078] | 0.0001 | supported |
| H4 no individual-accuracy harm (opportunity region)1 | -0.070 | 0.005 | [-0.080, -0.060] | 2.29e-23 | supported |
| H5 focused beats broad | -0.011 | 0.003 | [-0.017, -0.005] | 0.0001 | supported |
| Calibration control (normal cells; gate, not hypothesis) | 0.016 | 0.007 | [0.002, 0.030] | - | pass (0/40 wins) |
| 1 H4 is a one-sided non-inferiority test at margin 0.0488; its one-sided upper 95% bound is -0.0614. | |||||
| Source: primary rows of confirmatory_model_rows.tsv. Log loss-ratio scale; negative favors the flexible pipeline. One Holm family over five p-values, alpha = 0.05. Satterthwaite df; 2,400 frozen replicate rows (H4: 800; H1: 1,600; H5: 1,600). | |||||
All five primary hypotheses are supported under the locked specification, with the smallest Holm-adjusted p-value at \(1.2 \times 10^{-131}\) (H2) and the largest at \(0.000\) (tied for H3 and H5 by Holm’s step ordering). The calibration control passed with 0 wins among 40 normal cells. Effect sizes on the ratio scale: H1 corresponds to a loss ratio of .80 at the design center; H2 to a further multiplication by .89 per doubling of N; H4’s point estimate corresponds to a 7 percent individual-accuracy improvement over the opportunity region, with the seven flagged cells of Table 15.3 qualifying its scope; H5 to a 1 percent focused-over-broad edge whose magnitude is specification-dependent (Section 14.5).
| Condition | MSEL ratio | 95% CI | Label |
|---|---|---|---|
| 2PL, N = 500, reliability 0.5, bimodal | 1.138 | [1.115, 1.162] | block |
| 2PL, N = 200, reliability 0.5, bimodal | 1.122 | [1.095, 1.149] | block |
| Rasch, N = 200, reliability 0.5, bimodal | 1.094 | [1.048, 1.136] | caution |
| 2PL, N = 200, reliability 0.6, bimodal | 1.092 | [1.070, 1.116] | caution |
| Rasch, N = 500, reliability 0.5, bimodal | 1.076 | [1.059, 1.093] | caution |
| 2PL, N = 500, reliability 0.6, bimodal | 1.072 | [1.059, 1.084] | caution |
| Rasch, N = 200, reliability 0.6, bimodal | 1.052 | [1.030, 1.077] | caution |
| Source: generated F$safety$fired rows. The 7 opportunity-region cells whose DP-versus-Gaussian MSEL ratio exceeds the 1.05 pass boundary; every interval excludes 1. Addendum 006 is a post-outcome reporting policy: the full block anchors the confirmatory H4 record, while later synthesis may use a concise warning and cross-reference. | |||
15.3 The secondary family
| Hypothesis | Estimate | SE | 95% CI | p (Holm, secondary) | Decision |
|---|---|---|---|---|---|
| H6 reliability slope of the log ratio (non-normal, equal shape weight) | -0.514 | 0.037 | [-0.587, -0.441] | 4.38e-41 | supported |
| H7 summary-by-goal dissociation | 0.281 | 0.031 | [0.210, 0.351] | 8.57e-06 | supported |
| H6 diagnostic: the same slope within normal cells | -0.029 | 0.052 | [-0.131, 0.074] | 0.5831 (raw) | not evaluated (diagnostic) |
| Registered by Addendum 003 after the P-3 pilot and before any production fit; own Holm family of exactly two, no alpha shared with the primary family. Amendment 001 specifies H7's share-only floor and form_family clustering. | |||||
Both secondary hypotheses are supported in their own two-member Holm family. H6’s slope, -0.514 per unit of achieved reliability, describes moderation along the realized ladder, where reliability, item count and finite-test information move together; it is not a reliability-only effect. Its decomposition by shape in Section 10.2 is a post-registration exploratory analysis, and its normal-cell diagnostic is flat as required. H7’s dissociation, +0.281 log units, is registered; the broader lever analysis of Chapter 9 is a post-registration synthesis.
15.4 Every estimate under every registered specification
| Hypothesis | Primary (locked) | Equal condition weight | Strict-pass only | CR2 cluster-robust |
|---|---|---|---|---|
| H1 | -0.221 [-0.246, -0.195] | -0.221 [-0.246, -0.195] | -0.232 [-0.249, -0.215] | -0.221 [-0.285, -0.156] |
| H2 | -0.114 [-0.123, -0.106] | -0.114 [-0.123, -0.106] | -0.116 [-0.124, -0.108] | -0.114 [-0.123, -0.106] |
| H3 | -0.127 [-0.175, -0.078] | -0.127 [-0.175, -0.078] | -0.121 [-0.174, -0.069] | -0.127 [-0.158, -0.095] |
| H4 | -0.070 [-0.080, -0.060] | -0.070 [-0.080, -0.060] | -0.069 [-0.078, -0.060] | -0.070 [-0.081, -0.059] |
| H5 | -0.011 [-0.017, -0.005] | -0.011 [-0.017, -0.005] | -0.003 [-0.006, 0.000] | -0.011 [-0.015, -0.007] |
| H6 | -0.514 [-0.587, -0.441] | -0.514 [-0.587, -0.441] | -0.486 [-0.558, -0.414] | -0.514 [-0.604, -0.424] |
| H7 | 0.281 [0.210, 0.351] | 0.281 [0.210, 0.351] | 0.286 [0.218, 0.353] | 0.281 [0.242, 0.320] |
| Entries are estimates with 95% intervals from confirmatory_model_rows.tsv and secondary_confirmatory_rows.tsv on the registered log scale. CR2 intervals use 3.7-4.0 degrees of freedom and issue no family verdict by convention. | ||||
The grid contains 28 estimates; its reading is given in Section 14.5. For completeness of the confirmatory record, all six adequacy checks returned claim_wording_step = full, so no registered confirmatory statement required a wording downgrade. This status does not assign a wording tier to later exploratory claims.
15.5 Interpretation boundaries set by the registration
The registered record supports the following confirmatory claim classes. H1, H2, H3 and H5 support their stated distributional-recovery claims within their registered domains. H6 supports moderation along the realized reliability/item-count/information ladder, not a reliability-only causal effect. H4 supports mean non-inferiority over the opportunity region, qualified by the seven bimodal low-reliability cells displayed here and in Section 12.1. H7 supports the registered summary-by-goal dissociation.
The book is not otherwise confined to those claim classes. The evidence map, quantile and tail families, lever decompositions, shape-specific H6 slopes and other post-registration analyses may be used for descriptive explanation and exploratory synthesis. Their boundary is provenance: they do not acquire Holm protection, alter a registered verdict or license unqualified extrapolation beyond the simulated grid. Exploratory status is marked at the entry to a section or where a statement could reasonably be mistaken for a confirmatory one, rather than repeated at every use. Tail ratios remain especially fragile because of their zeros and floor incidence, and the evidence map’s harm-flag label remains a residual category rather than a harm verdict.