17 Discussion
We have reported a preregistered simulation study of two remedies for the default practice of Bayesian IRT reporting, flexible Dirichlet process mixture priors and goal-specific posterior summaries, crossed with each other and with the reliability of the measurement. The design comprised 120 conditions and 7,200 fits; the five primary and two secondary hypotheses were supported under the locked specification and its registered sensitivities, and the negative control was quiet. H4 passed on average, while seven bimodal reliability-0.5/0.6 cells had 5 to 14 percent higher MSEL, ratios 1.052 to 1.138 with every interval excluding 1; Section 12.1 gives the complete cell-level table and discloses Addendum 006’s post-outcome timing. This reporting safeguard does not constrain the book’s explicitly exploratory synthesis.
That synthesis is conditional on the implemented scope: dichotomous, unidimensional Rasch/2PL models; normal, skewed and bimodal shapes; N = 50–500; five joint length-plus-\(c^\star\) ladder tiers; and one item-bank template per family. It is not a universal ordering of priors, summaries or measurement designs.
The role of the realized information ladder. The study’s organizing claim was that the targeted information regime helps determine the effective combination of prior and summary, and the results give that claim a specific form. The implemented coefficient co-varies with the shrinkage the summaries address and with the likelihood resolution available to the flexible prior; the association is steep for bimodality and milder for skew. The top two tiers are where the two studied levers most often compound. Among the methods and conditions studied, estimator changes did not replace weak measurement information, and rank loss was empirically dominated by position on the same joint ladder. These are results along a length-plus-\(c^\star\) design, not causal effects of a scalar reliability intervention.
The role of the posterior summary. Against the field’s emphasis on prior flexibility, the summary was the larger lever in this design by a factor of several on every distributional family, and Gaussian + GR had lower condition-mean loss than focused-DP + PM in 100 of 120 cells. We regard this as the finding with the most immediate practical content because GR is computed by post-processing existing draws, whereas changing the prior requires another fit. It also reframes what a flexible prior is for: with a goal-mismatched summary its contribution is shrunk away before reporting, so the prior question is meaningful mainly for pipelines that have already adopted goal-matched summaries.
Where the flexible prior helped in this grid. The earlier parent manuscript reported a comprehensive simulation, including model-by-summary winner maps and case studies, rather than merely proposing a pilot conjecture. It supplied earlier empirical evidence that a Gaussian model with the right summary often suffices below reliability 0.7, whereas the DP model becomes most useful when high information, non-normal truth and a matching summary occur together. The present study re-examines those findings in a separately frozen, replicate-paired production design with registered contrasts and adds two refinements: an exploratory shape decomposition concentrates the reliability gradient in the bimodal cells (at N = 500, skew favors the flexible prior across all five simulated tiers), and in the mostly higher-information exception region the prior lever overtakes the summary lever. That region also contains a few high-reliability cells with N below 200 and 2PL skew cells at reliability 0.6, so it is not one rectangular corner. The recommendation there is usually the flexible prior with GR, while two paired-geometric exceptions select CB instead.
Limitations, in order of bite. First, the five tiers were constructed jointly. Nested test length supplies the coarse ladder, but a per-model-by-form-by-tier-by-shape global discrimination multiplier then refines the target; the frozen values span 0.794–1.090. Achieved reliability and log item count correlate above 0.98 in the estimation rows, but neither is a pure intervention, and the scalar match does not equalize full information curves across shapes. H6 therefore describes moderation along this realized ladder. It does not identify separate effects of length, discrimination or population geometry. Second, the generating conditions are stylized: three shapes, one item-bank template per family, no guessing, missingness, multidimensionality or person misfit, and dichotomous items only. Third, the tail-classification family is uninformative by construction at this design’s scale, its floor binding for 60 to 85 percent of comparators, so the study supports no tail claims; a future design should scale cutoffs to realized spread or report tail shares with binomial error. Fourth, point summaries are the study’s whole subject; interval estimates, coverage, and decision procedures built on full posteriors are outside it. Fifth, the magnitude of the smallest registered effect (H5) is specification-dependent, and one registered record (Addendum 006) was written after outcomes were known. Its full disclosure is retained with the direct H4 record, while summaries and discussion preserve the seven-cell warning and cross-reference that record rather than reproducing the full table at every mention.
What a next study should do. The agenda follows from the limitations. A factorial design that independently crosses nested length, global discrimination scale and difficulty geometry would separate their contributions and compare cells with the same scalar \(\bar\rho_J\) but different information curves. That experiment differs from the present shape-specific \(c^\star\) refinement, which was chosen to hit targets rather than to identify a discrimination effect. An empirically calibrated item-bank family would test the template’s externality; and extending the summary roster to interval and decision summaries would connect this design to the uses of posteriors the present study deliberately excluded. The machinery is built for such extensions: the design freeze, the ladder calibration, the paired replication structure and the claim-gating apparatus are all indifferent to which factors they cross.
Concluding remark. The study set out to learn when a flexible prior helps, and its most durable lesson is about the question’s frame: the useful unit of advice is not a model but a combination of model, summary and instrument. Within this design’s scope the results form a descriptive lookup among the studied combinations; more often than not, its largest loss differences involve the summary and position on the joint ladder rather than the prior.