| claim_id | frozen_text | corrected_text | reason |
|---|---|---|---|
| C-06 | … changes about 18% of the top decile’s membership (top-10 overlap 0.818 against 0.923 under skew). | … changes about 10% of the top decile’s membership (top-10 overlap 0.818 against 0.923 under skew, i.e. 10% of the selected group changes against 4%). | Arithmetic error converting Jaccard overlap to a share of the selection. For equal-sized sets of size k, |intersection| = 2kJ/(1+J), so J=0.818 leaves 90% of the top decile shared and 10% changed; 18% is the share of the UNION lying in exactly one method’s selection. The evidence (overlaps 0.818 and 0.923) and the claim’s direction are unchanged; only the reported percentage was wrong, by nearly a factor of two. |
| C-07 | The elicitation setting is close to irrelevant for reported outputs: DP-Focused against DP-Broad moves individual scores a median 0.005-0.025 SD, with SD ratios of 0.996-0.999. | Negligible for scores, distribution and tails but NOT for rankings: 4 of 78 material on scores and 0 of 78 on cut scores, but 45 of 78 on rankings, median top-decile overlap 0.887 (about 6% of the selected group changes) – the same median the prior-family contrast produces. | The claim generalised a null across all four output families from evidence on two of them. Its own evidence table (P3-T7) carried the top-decile overlap column showing 0.882-0.887 throughout, and the wording did not read it. No number changed; the scope of the null did. Reported as a null on rankings, the claim would have told a practitioner selecting a top decile that a consequential choice was free. |
Appendix D — Deviations, corrigenda, and the release record
This appendix preserves the frozen v1 claim record and states the changes that make each subsequent release substantively different from the last. Neither v3 nor v4 claims that its changes are merely presentational.
D.1 The v1 correction record
The frozen register also contains one qualification, two widenings, and two exploratory observations. E-01 is the clump/band-edge cut mechanism. E-02 is the structured sub-material normal-control profile and is labeled exploratory in Chapter 10; neither is promoted to a settled finding.
D.2 V3 load-bearing corrections
- Shape authority. The legacy dip statistic used fixed native-scale jitter. Because the stored null-fit quadrature supports and weights were unavailable, v3 marked all 26 legacy classes unvalidated and set recommendation eligibility to false, supplying a corrected standardized statistic for a future full refit. V4 ran that refit and closed this item; see the next section.
- Reliability placement. V3 reports continuous \bar\rho, nearest tier, gap, and support. C12–Rasch and C4–2PL are below .5-tier nearest support and are not matched.
- Verdict provenance. The case-to-simulation join is rebuilt from sim-v3, labeled unpreregistered and exploratory, and reduced to 19 in-support rows. V3 made no case recommendation from them. V4’s validated shape coordinate reopened that gate: the same 19 rows now carry case recommendations; see the next section.
- Safety timing. Addendum 006 is identified as post-outcome. Its H4 interpretation is not rewritten as a blanket withholding policy.
- Estimands. Factorial lever shares, condition-mean ratios, and paired-geometric ratios are named separately. A 12% identity tolerance was replaced by an exact current-authority check.
- Summary theory. CB is described as a marginal-variance, posterior-independence specialization; GR targets the expected finite-sample EDF and expected ranks.
- GR mechanism. Exact ties do not explain the instability. V3’s audit uses the common Gaussian-PM cell SD, identifies 11 failures, and attributes them to dense near-rank reassignment of a stable quantile ensemble.
- Comparator framing. The Gaussian + PM ratio to a selected winner is reported together with the Gaussian + GR comparator ratio.
D.3 V4: the shape refit, and what it settled
V3 left one Critical finding mitigated rather than fixed, with an explicit reopening condition: a standardized observed-and-null refit. V4 performed it.
What was run. 26 observed empirical-histogram refits and 5,158 retained null replicates of 5,200 attempted, reproducing the v1 characterization exactly (same samples, convergence ladder and seed tree) and changing only the dip estimator, which now standardizes the weighted support before jittering. The same estimator was applied to the observed densities, to every null replicate and to the simulation’s reference conditions, so the three quantities the audit found incommensurable are now computed one way. The fitted supports and weights that v1 discarded are retained this time, under data/shape-refit/, so no future correction of this class needs a refit.
What it showed. The scale-invariant statistics reproduce the frozen values to better than 0.00018 on 24 of 26 cells, which is what establishes that the refit is faithful; the two exceptions are numerically borderline density fits (C12–Rasch and C9–Rasch). 1 label of 26 changes, and the old estimator reproduces that same change, so it is refit variation rather than a consequence of standardizing. The simulation’s reference dip is 0.054 unjittered against 0.054 standardized, so the half-condition gate’s jitter asymmetry was immaterial. The dip’s correlation with the fitted latent spread falls only from 0.59 to 0.54: standardization removes an incomparability, not the real association between bimodality and variance.
What changed in the book. The shape coordinate is validated, so Chapter 15 is a matched-and-judged chapter again rather than a sensitivity join, and 19 rows are recommendation-eligible. Everything the v3 edition reported by shape class in Part IV stands unchanged, because the labels did. The v3 caveat that the labels were unvalidated is withdrawn; the v3 caveats about the reliability tier gap, the paired seed reference and the gallery’s normal overlay all stand. V4 also documents the classifier’s mild-* and normal-detectability branches in Chapter 5, which no earlier edition described.
What did not change. The reliability coordinate. A case’s \bar\rho is a fitted-model quantity and the grid’s tier is a design-time quantity; they are commensurable in definition, not identical in construction, and every matched row is reported with its gap (median 0.033, maximum 0.049). Closing that gap would require a simulation designed around fitted coordinates, which is a study, not a repair.
D.4 V3 numerical and rendering corrections
The release also corrects the Chapter 12 percentage and churn conversions; tail outlier attribution and band shares; the GR denominator mismatch; null-calibration and coordinate-map captions; score-plot clipping; corpus funnel arithmetic; total computation; and stale/blank freeze output. Exact values and validation checks are recorded in the phase logs stored with the external review.
D.5 Release-status rule
Every result is assigned one of three roles: current authority, exploratory sensitivity, or legacy audit artifact. A frozen historical value may be displayed for provenance, but it may not silently override the current metric definition. The clean build runs semantic checks before rendering and records input hashes in data/derived/input-provenance.csv; the standardized shape manifest is manifest/shape-classification.csv, and the per-label comparison against the legacy screen is data/shape-refit/shape-audit.csv.