4  The corpus and the screen

The Item Response Warehouse (Domingue et al., 2025) harmonizes item response datasets from published research into a single long format, one row per person-item response, and it is what makes a study like this one possible at all: thirteen real assessments, fitted under identical machinery, with their provenance on record. This chapter describes how 996 catalogued units became thirteen cases, and what the corpus itself reveals about where real tests live on the simulation’s grid.

4.1 The funnel

A funnel of horizontal bars from 996 catalogued units through 889, 572, 138 and 29 to the 13-case portfolio.
Figure 4.1: From nearly a thousand catalogued warehouse units to thirteen cases: most of the loss is structural, not selective. Bars trace the corpus funnel: units catalogued in the Item Response Warehouse; the subset whose reliability the project’s first corpus paper could measure; the subset with measured shape; the dichotomous units (the DPMirt item models fit only these, which alone removes about three quarters of the shape corpus); the 29 that pass all seven eligibility rules; and the 13-case portfolio (orange), which adds three small whole-corpus cases and two low-reliability recruits from outside the automatic screen. A separate 54-unit small-sample frame, not shown, is used only to demonstrate why small datasets cannot be shape-classified. Counts from the frozen corpus tables.

Most of the attrition in Figure 4.1 is structural. Two earlier corpus papers of this project supply the screening measurements: a reliability audit covering 889 units (Lee, 2026b) and a latent-shape audit covering 572 (Lee, 2026a). The single largest cut is dichotomy: the DPMirt item models fit binary responses only, which removes roughly three quarters of the shape corpus and leaves 138 units. Seven eligibility rules then screen for a single administration, at least a thousand respondents, 12 to 60 items, approximate unidimensionality, a converged empirical-histogram fit, and a plausible response-matrix density, leaving 29 units eligible for the characterized tier. The portfolio adds three small datasets fitted whole (their own tier, with the screen relaxed) and two deliberate low-reliability recruits, for reasons the next section makes concrete.

Two properties of the corpus deserve the reader’s suspicion before any result. The warehouse over-represents shareable data: internet samples, volunteer panels, and instruments whose licences permit deposit, while high-stakes operational testing is nearly absent (one case here, the Brazilian ENEM, is the exception). And it over-represents clean data, because depositors curate. Both biases limit coverage claims, neither touches internal comparisons: every contrast in Part IV is within-dataset, so the samples’ unrepresentativeness cannot manufacture a difference between two methods applied to the same people.

4.2 Where real tests live

The corpus concentration matters to anyone who wants to use the simulation’s grid, because the grid is uniform and reality is not. Figure 4.2 places the screened pool on the reliability axis: of 60 screenable units, 28 sit in the .75–.85 band and only 6 fall below .65, several of those pathological. The simulation spends two of its five tiers below that line. A reader should therefore expect the high-reliability rows of the simulation’s maps to do most of the work on real tests, and should treat its \bar\rho = .5 row as covering a rare and often broken corner of practice, a point the portfolio’s one inhabitant of that corner (C12) will make vivid.

Histogram of Rasch marginal reliability over the eligible pool, concentrated at .75 to .85, with the thirteen portfolio cases marked as ticks.
Figure 4.2: Real dichotomous tests concentrate at reliability .75 to .85; the simulation’s lowest tier is close to empty in the corpus. The histogram gives the Rasch marginal reliability \bar\rho of the 60 screenable units in the eligible Item Response Warehouse pool (bin width .05); dashed vertical lines mark the simulation’s five design tiers. Of the 60 units, 4 fall below \bar\rho = .55 and 2 in .55-.65, against 28 in .75-.85. Orange ticks locate the thirteen portfolio cases (Rasch coordinate); the portfolio deliberately over-samples the sparse low end. C12 at .43 lies below the simulation ladder’s nearest-tier support and therefore receives no matched-cell recommendation in this edition; C13 sits at .65. Corpus values from the v1 screening frame (60 of 64 units computable).

4.3 Licences and identification

About a third of the corpus carries no usable licence field. Catalogue status needs more precise language: 149 of 996 units (15%) are actually provisional_grade_b, while 538 (54%) carry the accepted-but- ambiguous status accepted_ambiguous_v1_1. The portfolio prefers core-status units but does not require them (two cases are provisional, and say so in their dossiers). Identification is the quieter hazard. For three of the thirteen cases the catalogue’s own descriptive fields were wrong or empty, one construct field held the title of an unrelated methods paper and an age field read “Non-human”, and the instruments had to be identified from source publications or shipped documentation. The dossiers in Appendix A record, for every case, what the identification rests on. None of this is a complaint about the warehouse, whose harmonization is what made the audit possible; it is a caution that a corpus row is a claim about provenance, not a fact, until it has been checked against the instrument it names.

Domingue, B. W., Braginsky, M., Caffrey-Maffei, L., Gilbert, J. B., Kanopka, K., Kapoor, R., Lee, H., Liu, Y., Nadela, S., Pan, G., Zhang, L., Zhang, S., & Frank, M. C. (2025). An introduction to the item response warehouse (IRW): A resource for enhancing data usage in psychometrics. Behavior Research Methods, 57(10), 276. https://doi.org/10.3758/s13428-025-02796-y
Lee, J. (2026a). How common are estimated latent-distribution departures from normality? Evidence from 504 item-response data sets. https://doi.org/10.48550/arXiv.2608.06817
Lee, J. (2026b). How reliable are psychological measurements? The distribution of marginal reliability across 889 item-response datasets. https://doi.org/10.48550/arXiv.2608.06806