Appendix D — Version history and lessons
This appendix uses V1, V2 and V3 for the three successive simulation-study builds. The current Quarto book v3 is a reporting remediation of the V3 production study; it is not a fourth simulation wave and does not replace the frozen confirmatory estimates. Keeping the two version sequences distinct is necessary because earlier builds informed design choices without acquiring confirmatory status.
D.1 V1: exploratory results prompted a governed design
The first rebuild used a wider, thinner grid without a locked hypothesis family. Its post-run review did not support the original general headline that the flexible prior wins. Because the estimands and multiplicity procedure had not been fixed prospectively, V1 is not used as confirmatory evidence here. Its main contribution was procedural: V3 uses a locked model-based hypothesis family, a single Holm adjustment for H1–H5, and a refusal rule when a registered contrast cannot be computed as specified.
D.2 V2: useful preliminary evidence from an uncertified design
The second rebuild framed the comparisons that V3 later tested, but its reliability ladder could not be certified. Production used a ladder whose own record was blocked, while a gate had accepted evidence with the same label from a different candidate artifact. Achieved reliability overshot most targets and the top two tiers were insufficiently separated. A later audit also found that the generating and fitted reporting scales had been named as if they matched without an explicit map. V2 results are therefore never treated as confirmatory evidence. They were nevertheless used transparently as earlier simulation evidence to motivate the V3 questions and the H4 opportunity region.
Other measured V2 limitations also shaped V3: the DP arms had lower Monte Carlo precision than the Gaussian arm, a declared single-chain audit was not run, and fixed-sequence gatekeeping made inference depend on test order. V3 addressed those limitations through arm-specific production budgets (Section 14.1), four-chain fitting, and model-based multiplicity control.
D.3 V3: the five specified fixes
The governed rebuild was specified around five fixes recorded before production:
- calibrate and certify the achieved-reliability ladder before fitting;
- replace fixed-sequence gatekeeping with a preregistered coefficient family and Holm adjustment above a non-claim-generating descriptive cell map;
- run four dispersed-start chains for every production fit and make chain independence a blocking check;
- size Monte Carlo replication from measured within-cell variance in V2 rather than from a default; and
- test the fit-engine execution contract—seeds, generator identity, custom distributions and process-level parallelism—under the same session construction used in production.
Two architectural requirements accompanied them: pre-generate and audit the entire response store before fitting, and preserve every fit bundle so later questions can be answered by re-reading rather than refitting. The execution incidents in Section A.3 do not alter that list; they show how the identity and execution checks operated during the wave.
D.4 The V3 registration sequence
The plan was locked on 2026-07-24. P-2 had already measured mixing, throughput and storage. Addenda 001 and 002 were then recorded on July 26–27, before P-3 and production: they fixed the H1 reference grid and the symmetric replicate-level floor, respectively, while excluding condition-level ratios from that floor. P-3 subsequently exercised the full analysis at one-tenth scale. Addendum 003 followed P-3 but preceded every production fit; it disclosed the pilot prompt and added H6 and H7 as a separate Holm pair. Its Amendment 001 specified H7’s share-only floor and form_family clustering.
Addendum 004 was recorded after 1,798 production fits existed but before a production analyzer or outcome. It added an optional direction-blind replication extension; it did not create the wording ladder already present in the lock. Addendum 005 was post-fit and pre-analysis: it repaired the common reporting-scale map and restricted H6 to the realized reliability/item-count/information ladder. Addendum 006 alone was post-outcome. It governs proportionate reporting of the seven H4 safety cells; it neither re-tests H4 nor restricts unrelated exploratory synthesis. The source-backed record is consolidated in Chapter 15.
D.5 The reporting remediation in book v3
External review of the preceding Quarto edition independently reproduced the confirmatory estimates but identified reporting defects: the reliability coefficient and joint length–discrimination calibration were overstated, parts of the theory exceeded their sources, the addendum chronology was compressed incorrectly, some findings were hard-typed, and several captions or scope statements did not match their generating rows. This book v3 repairs those descriptions, provenance paths and figure semantics while retaining the frozen row-sets and registered verdicts. It also keeps two layers explicit: the confirmatory record is fixed, whereas clearly identified post-registration analysis remains available for free descriptive and exploratory synthesis.
D.6 The pattern worth carrying
Across builds, the recurring failure was a check on a name when the scientific requirement concerned content or a relationship: the exact calibrated artifact, the reporting-scale map, the realized throughput, or the agreement between plotted values and captions. The reusable response is to bind those relationships to checkable artifacts—hashes that travel with rows, replay calculations, source-backed facts, copy-equality checks, and semantic figure QA. These checks add modest runtime and make the scope of each result auditable.