flowchart TD
A["What is the report for?"] --> B["Individual scores"]
A --> C["Ranking / selection"]
A --> D["Distribution / shares"]
B --> B1["Posterior mean.<br/>Gaussian is the defensible default<br/>among the studied arms."]
C --> C1["Estimator differences were small in this grid.<br/>Prioritize measurement information."]
D --> E{"Position on the study's<br/>information ladder?"}
E -->|"up to ~0.7"| D1["Gaussian + GR is a defensible default.<br/>Consider DP for skew at large N."]
E -->|"0.8 - 0.9"| F{"Non-normality plausible?"}
F -->|"yes"| D2["DP (focused) + GR performed best.<br/>Check the individual-accuracy tradeoff<br/>if scores are also reported."]
F -->|"no"| D3["Gaussian + GR among the studied arms."]
16 Guidance within the studied design
This chapter restates the results as decisions, for an analyst holding one dataset from one test and owing one report. The evidence behind every rule is cited to its chapter; the rules themselves are conditional on the study’s scope: dichotomous unidimensional data, normal, skewed and bimodal latent shapes, N = 50–500, five ladder tiers from 0.5 to 0.9, and one item-bank template per Rasch/2PL family. The operational coefficient varies along a joint nested-length plus cell-specific-\(c^\star\) ladder, so these recommendations do not isolate a scalar reliability effect. Chapter 17 gives the full delimitations.
This chapter is an exploratory decision synthesis, not a preregistered decision rule. It translates the registered results and the book’s post-registration analyses into conditional defaults for settings resembling the simulation; claims about other item banks, latent shapes or operating constraints require new evidence.
16.1 The decision sequence
16.2 The rules and their tradeoffs
The role of the summary. Match it to the goal first; among the studied choices this produces the largest observed distributional-loss change and requires only post-processing. Using the posterior mean rather than GR for distributional reporting raises geometric KS loss by 31 to 73 percent across the displayed reliability bands (more on the quantile scale), while computing GR from existing draws costs seconds (Chapter 9). If one estimate set must serve both scores and shares, the choice is a genuine tradeoff: GR’s tier-median individual MSEL is about 10 to 20 percent above PM in this grid, while GR has lower distributional loss. The 240 displayed PM-versus-GR condition-mean point estimates all occupy that dissociation quadrant; H7 provides the pooled registered test (Section 9.3).
The role of the prior. Within this grid, the flexible prior is most useful when three conditions hold together: distributional goals, the studied non-normal shapes, and ladder tier 0.8 or above (earlier for skew when N is large). In the most favorable cells it reduces distributional loss by up to half; below ladder tier 0.7 its bimodal advantage is generally smaller, and seven bimodal low-tier opportunity-region cells show 5 to 14 percent higher individual MSEL, ratios 1.052 to 1.138 with every interval excluding 1 (Chapter 12). Under defensible normality in this grid it adds little. The calibrated concentration prior is a defensible default, but the primary H5 advantage over the diffuse arm is only about one percent (log ratio -0.01099). Under strict-pass-only filtering it attenuates to -0.00284 (two-sided 95% CI [-0.00610, 0.00042]; registered one-sided Holm \(p = 0.0399\)), so this is not evidence of a practically large focused-over-broad advantage (Section 8.5).
The role of reliability. For ranks it is the dominant lever observed among the choices studied here (Chapter 13); for the other studied goals it is the joint design axis that co-varies with both the need for repairs (shrinkage grows as the ladder falls) and observed shape recovery (which attenuates in shorter, lower-tier forms). An assessment program that will report distributions should treat the reliability tier as part of the reporting design, not merely as a psychometric summary. Between tiers 0.7 and 0.9 the observed ordering shifts toward the flexible prior under the studied non-normal shapes, with loss ratios reaching about one half. Because length and \(c^\star\) also move, this is an association along the implemented ladder rather than a reliability-only effect.
The compute bill. The DP arms cost roughly 1.5 times the Gaussian arm’s iterations under Rasch, and the 2PL family costs several times the Rasch family at matched reliability (Section 11.3); at this study’s scale the whole wave was 147.7 host-hours for 7,200 fits. For a single operational dataset the marginal cost of applying these workflows may be minutes, depending on software and hardware; within this implementation the summary was not the expensive component.
16.3 What we would tell a colleague in one paragraph
Decide what the report is for before deciding what to fit. If it is for individuals, the familiar Gaussian model with posterior means is a defensible default among the studied arms. If it is for an ordering, prioritize measurement information because estimator differences were small in this grid. If it is for a distribution, compute the triple-goal summary for the prior families studied here; then, in settings resembling the higher-tier non-normal cells, fit the DP mixture as well and expect lower distributional loss. If individual scores are also reported, retain the registered disclosure that seven bimodal low-tier cells had 5 to 14 percent higher individual MSEL, ratios 1.052 to 1.138 with every interval excluding 1. These are exploratory defaults for comparable settings, not rules for other item banks or measurement designs.