6  Machinery and the locked plan

This chapter covers the apparatus between design and results: the computational pipeline and its budgets, the pilot ladder that qualified it, and the preregistered analysis plan, with the registration history stated plainly. Readers headed for the results need three things from it: what a fit is, what was specified before production outcomes existed, and which specifications were amended when.

6.1 The pipeline

flowchart LR
  A["Design freeze<br/>120 conditions, seeds"] --> B["Response store<br/>2,400 datasets, frozen"]
  B --> C["Fitting wave<br/>3 arms x 2,400 = 7,200 fits<br/>4 chains each"]
  C --> D["Extraction<br/>PM, CB, GR per fit<br/>common reporting scale"]
  D --> E["Losses<br/>16 rows per fit<br/>115,200 rows"]
  E --> F["Aggregation<br/>form, condition, contrast<br/>bootstrap B = 10,000"]
  F --> G["Confirmatory models<br/>H1-H7, one Holm family + pair"]

The production pipeline. Datasets are frozen before any fitting; every fit is summarized three ways and scored on five losses; all analysis rows descend from one frozen store with content-addressed identities.

Fitting uses MCMC via NIMBLE (Valpine et al., 2017) with the project’s DPM implementation, four chains per fit. The response store is generated once, frozen, and content-addressed; no dataset was ever regenerated, and every downstream row carries hashes tying it to its dataset and fit bundle. The analysis stage is resumable per fit and was verified by replaying a sample of stored loss values from stored draws before aggregation (36 of 36 replays matched; Appendix A).

6.2 Budgets and diagnostics

Sampler budgets are the second disclosed family asymmetry. Rasch fits ran 8,000 iterations per chain for the Gaussian arm and 12,000 for the DP arms (halves discarded as burn-in, thinning 2). A lean 2PL pilot failed its mixing checks in the discrimination block, and the 2PL budgets were raised, once and before production, to 40,000 (Gaussian) and 36,000 (DP) with thinning raised in proportion, so retained draws per chain are 2,000 and 4,000 everywhere. Diagnostics are computed over the full reporting parameter set as rank-normalized split R-hat with bulk and tail effective sample sizes; the preregistered policy gates claim wording on diagnostics but never drops a completed fit from analysis. Thresholds: warning at R-hat 1.01, hard failure at 1.05, bulk ESS floor 400, person-level minimum 100. Realized outcomes, including the zero hard failures, are in Section 14.1.

6.3 The pilot ladder

Production was preceded by four qualification rungs. P-0 exercised the engine end to end on toy data; P-1 rehearsed the data factory and froze the store format; P-2 measured throughput, memory and mixing at scale and is the rung that exposed the 2PL budget deficit; P-3 ran the full pipeline at one-tenth scale, producing every table and figure the production run would produce, watermarked as non-citable, so that the analysis code encountered pilot data before production fitting began. The P-3 outputs remain permanently non-citable, and none of their numbers appears in this book.

6.4 The locked plan

The plan was locked before the production wave. Its estimand is the replicate-level paired log loss ratio; its confirmatory model is a linear mixed model on the 2,400 replicate rows with fixed effects for shape, log2 sample size, centered achieved reliability and model family (with the registered interactions) and a random intercept for the ten form families; its family is one Holm procedure over exactly five p-values at \(\alpha = 0.05\). In words:

  1. H1, the flexible prior improves KS distribution recovery under non-normality (one-sided);
  2. H2, the improvement grows with sample size (one-sided slope);
  3. H3, the improvement’s shape profile differs between model families, Rasch pairing with bimodal and 2PL with skew (two-sided difference in differences);
  4. H4, over the opportunity region (non-normal, N of 200 or 500) the flexible pipeline’s mean individual-loss increase stays below 5 percent (one-sided non-inferiority at margin log 1.05); and
  5. H5, the calibrated concentration prior beats the diffuse one (one-sided).

Two secondary hypotheses were added by Addendum 003 after the P-3 pilot but before any production fit existed, with their own two-member Holm family: H6, the advantage depends on achieved reliability within the non-normal domain, with a within-normal diagnostic slope required alongside; and H7, the summary that minimizes loss depends on the goal, expressed as the PM-versus-GR dissociation across loss families. The addendum disclosed that both hypotheses were prompted by P-3 observations; those pilot results remain non-citable and do not enter the production estimates. Amendment 001, still before any production fit, specified that H7 uses the share component of the floor rather than its absolute component and clusters its primary standard error on form_family, with the coarser form_rep clustering retained as a sensitivity. The primary H1–H5 family was unchanged.

The opportunity region for H4 was restricted by design because the previous study version had already shown individual-accuracy harm for skewed populations at N = 50; testing non-inferiority where inferiority was expected would not have addressed the intended uncertainty, and the small-N cells are reported descriptively instead (Chapter 12).

6.5 Gates, controls, and the amendment record

Three mechanisms police the plan. The calibration control turns the normal cells into a blocking alarm: if the flexible prior wins where nothing should be won, at a rate above 0.05, every claim in the study is blocked; its realized rate was zero (Section 7.1). The MSEL screen’s block label fired. Addendum 006, written after the outcomes were known, resolved its reporting consequence: the full seven-cell safety block anchors the confirmatory H4 record, while later synthesis may preserve the warning concisely and point back to that anchor rather than reproduce the entire table at every mention (Section 12.1, Chapter 15). This is a reporting policy, not a re-test or a restriction on distributional or exploratory analyses.

The replication-adequacy rule in Addendum 004 was also pre-outcome, but not pre-fit: 1,798 production fits already existed when it was logged. There was no production analyzer and no production outcome. It added an optional, direction-blind replication-extension route; the wording ladder was already in the locked plan. The trigger did not fire. The complete record is Table 15.1 in Chapter 15. Addenda 001 and 002 were logged after P-2 but before P-3 and production: the first fixed the H1 reference grid, and the second applied the registered replicate-level loss-ratio floor symmetrically while leaving condition-level ratios outside its scope (Section 13.3). Addendum 005 was a post-fit, pre-analysis repair of the common reporting scale and also confined H6 to the realized reliability/item-count/information ladder (Section 5.3).

The distinction used in the rest of the book is functional rather than restrictive. The confirmatory record preserves the registered estimands, families, decisions, inclusion policy and timing. Post-registration synthesis is free to ask additional questions, refit descriptive models and combine results for explanation, provided it is not presented as a registered verdict and does not silently redefine one. Exploratory status is stated where it matters for interpretation and at the entry to an exploratory section; it need not be repeated mechanically in every sentence or caption.

Valpine, P. de, Turek, D., Paciorek, C., Anderson-Bergman, C., Temple Lang, D., & Bodik, R. (2017). Programming with models: Writing statistical algorithms for general model structures with NIMBLE. Journal of Computational and Graphical Statistics, 26(2), 403–413. https://doi.org/10.1080/10618600.2016.1172487