3  What a case study can and cannot settle

Two kinds of statement appear in this book, and they rest on different ground. Keeping them visibly apart is the book’s central discipline, so we state the rule once, early, and in full.

The first register is consequence, and it belongs to the cases. Without knowing any person’s true ability, a real dataset can still establish how much the reported output changes when the analyst’s choice changes: scores move, the reported distribution widens or changes family, a selected group gains and loses members, a share beyond a cut rises or falls. Statements in this register use a deliberately limited vocabulary (differs, moves, reorders, reclassifies, disperses) and never say that one method is better, closer, or more accurate on a case, because on a case no such thing is identifiable. Consequence is what Part IV reports.

The second register is direction, and it belongs to the simulation. Where truth is known by construction, “which combination recovers it” is a measurable question, and the companion volume measured it under a locked plan. Statements in this register appear only with their attribution, estimand, grid scope, and timing: the simulated cell’s ratio is, the evidence map labels, the post-outcome safety interpretation says. Part V transfers those verdicts to the 19 eligible case-cells through a join that remains exploratory in status, because it was assembled after both source volumes’ results.

The join between the registers is a lookup, and it is worth being precise about what the lookup does and does not assume. Candidate coordinates—sample size, fitted reliability, and a shape label—can select a simulation cell only when measurement and support checks pass. The v3 edition’s shape labels failed that validation gate; the standardized refit of Section 5.4 closes it, so the displayed join now rests on a validated shape coordinate and an in-support reliability placement. Nothing about the case outcomes validates a simulated direction, and a safety label is goal-specific rather than a generic reason to stop every analysis. Chapter 17 separates these failure modes.

3.1 Materiality, not distinguishability

A comparison needs a bar, and the wrong bar is statistical distinguishability. With five hundred respondents, MCMC noise is so small (measured, not assumed: the Monte Carlo standard error of a median movement is on the order of 0.001 to 0.003 SD here) that essentially every contrast in this book is distinguishable from zero, including ones no analyst should act on. The bar used instead is materiality: a pre-declared size, per output family, above which a difference is worth acting on. The thresholds (Chapter 7) were fixed before any model was fitted, and every count of “material” cells in this book is a count against them. The normal control cases will illustrate why the distinction matters: their movements are ten to forty times the sampler noise and still quiet by the materiality standard, which is the correct simultaneous reading of “the effect is real” and “the effect does not matter.”

3.2 The paired seed reference

One methodological device recurs so often that it needs a name here. Every one of the 78 production fits was re-run once with an identical specification and a different random seed. The pair gives a measured, not theoretical, reference for every comparison: a contrast between methods should be read against how much the same method differs in this A/B pair. The paired difference is large for one output family (a top decile churns 2.0 to 2.1% of its membership on a new seed; Chapter 12) and diagnostic for another (one summary fails to reproduce across seeds on short forms while the convergence diagnostics stay clean; Chapter 14). One pair per cell does not estimate the full seed-to-seed distribution, but it supplies information that a within-chain standard error does not.

3.3 Frozen artifacts

This edition keeps the frozen v1 coordinates, thresholds, claim register, production fits, and seed-B refits as historical inputs. It also creates current derived facts, corrects metric definitions, rebuilds the sim-v3 evidence join, and replaces the legacy shape authority with the standardized manifest of Section 5.4. Every upstream authority file used by the join is hashed in data/derived/input-provenance.csv. Frozen values may be shown for provenance, but they do not silently override a corrected current definition. The old and new records are reconciled in Appendix D.