23  Open Problems

This is the shortest chapter and it does the most to say what kind of document this is. In the first edition it closed the Rasch-centred theory spine by listing what that argument could not settle. In this edition it still closes that spine, but it does not end the book. It is the hinge to Part VIII, What Changes Under the 2PL, and to Part IX, which places the ledger beside realized evidence.

The hinge matters because the honest version of the argument in Chapter 21 — that the question changes when reliability is a design quantity — only carries weight if the questions it leaves open are stated as plainly as the ones it closes.

Each entry says three things: what is actually known, what would settle it, and whose problem it is. That last distinction matters. Some of these are places where this book had the tools and did not do the work. Some belong to the companion simulation study by the boundary this project set. And some are open in the field.

A note for the reader of this edition: the entries in the ledger below are preserved in their first-edition state, because rewriting an open-problem record after the fact would launder it. The surrounding navigation is updated. Readers asking how the item-model change affects the theory should continue through Chapter 24, Chapter 25, and Chapter 26. Readers asking what became of the ledger should continue through Chapter 27 and Chapter 28 to Chapter 29, which walks it item by item. Several entries are now settled by the companion volumes, some in a sharper form than the entry asked for, and the evidence has added entries of its own.

23.1 Problems this book could have closed and did not

23.1.1 The rank proposition outside its conditions

Proposition 20.1 is stated with conditions because the general claim is false. What is established is the leading-order statement — rank recovery depends on the model through the stability coefficient — plus a measurable residual under a matched constant-error Gaussian mapping and a Rasch computation on one grid of shapes.

Three things are unknown. Whether the residual’s behaviour is general: on the displayed grid the shape gap rises to a peak near \(\bar w = 0.915\) and then declines, and the decline to zero in the perfect-information limit is proved only for continuous working shapes. Whether the ordering result of Chapter 12 — that in a common-form Rasch model differential shrinkage cannot reorder — survives the three regimes Chapter 20 names as vulnerable: different test forms, incomplete designs, and item uncertainty carried into the posterior. And whether the matched-reliability construction is the right comparison at all, since it holds \(\bar w\) fixed while the shapes differ, which is one of several defensible matchings.

The first two are finite computations. They were not run.

23.1.2 How far the incompatibility theorem generalizes

Theorem 17.1 is an existence statement, and it is stated that way because an earlier draft asserted a universal one and was refuted (Chapter 17). The example is two independent non-degenerate overlapping normals; the degenerate coincident case where the actions agree is exhibited alongside it.

What is not known is the boundary. For which families, and under what separation of the posterior means relative to their spreads, do the WSEL and ISEL actions coincide? The question has a clean form — characterize the set of posterior ensembles on which the midpoint-quantile action equals the vector of posterior means — and it is not answered here. A useful partial answer would be the exchangeable equal-variance case, which is the one applied work most often has.

23.1.3 What a finite polytomous test identifies about \(G\)

Chapter 22 establishes that the partial credit model keeps sufficiency and the graded response model does not. It does not ask the semiparametric question. San Martín et al.’s result — that, under the anchored Rasch conditions, \(I+1\) displayed integral evaluations of \(G\) generate the observable law subject to their identities (Chapter 16) — has no counterpart here for a test of \(I\) items with \(m_i\) steps each.

The conjecture worth stating is that the count should depend on the total number of steps rather than the number of items, since that is what the sufficient statistic counts. It is a conjecture. Nothing in this book supports it beyond the analogy.

23.1.4 Location conventions and interpretable functionals of \(G\)

The earlier version of this section asked whether the location constraint changes the elicited prior on occupied-cluster count \(K\). It does not. The implemented Rasch model centres the item parameters and does not transform the DPM atoms; even if all atoms were translated by one draw-specific constant, their cluster memberships and occupied count would remain exactly unchanged.

A narrower question remains. Alternative location representatives, and the additional scale transformation required under a 2PL model, change the numerical locations and spreads used to describe \(G\). This book does not characterize how priors on decision-relevant functionals — such as tail mass beyond a fixed substantive cut, variance, or separation between mixture components — transform across those conventions, or which such summaries are invariant. The open object is therefore an interpretable functional of \(G\), not \(K\).

23.1.5 What the normal working model costs on short tests

Chapter 11 derives the closed-form posterior under a Gaussian working model and V8 verifies it against numerical integration. The working model is an approximation to a posterior whose error variance depends on \(\theta\), and the approximation degrades as the test shortens. The size of that error at the test lengths this work is aimed at — a dozen items or fewer — is not quantified. It is a one-dimensional numerical exercise and it was not done.

23.2 Problems the boundary assigns elsewhere

Two questions in this book’s neighbourhood belong to the companion simulation study, and saying so here is not evasion but the same boundary Chapter 21 draws.

Whether Lee et al.’s prior-by-information pattern transfers to the IRT design. Lee et al. show a conditional pattern in \(I\) and \(J\). Chapter 21 does not place assessment on the same numerical axis without a mapping. What happens under a full item-response likelihood as test length changes is a realized outcome.

Whether any of the theory’s preferences survive Monte Carlo error. Every statement in this book is about what is true of a model, not about what a finite simulation can distinguish.

23.3 Problems open in the field

23.3.1 Finite-\(N\) behaviour of the DPM posterior under weak identification

This is the deepest of them. Under the anchored Rasch conditions, a finite test’s observable law is generated by the \(I+1\) displayed integral evaluations of \(G\), subject to their identities, while guarantees for flexible deconvolution estimators are asymptotic. The literature read for this book does not give a finite-sample account of how posterior uncertainty in unidentified directions affects decisions based on those identified functionals.

Chapter 22 supplies a bounded comparison, not a solution. In restricted latent class models, grouped masses over finite equivalence classes can be identified when the relevant known-item or unknown-item structural conditions hold. Those masses are not partitions of continuous \(G\), and San Martín’s functionals already are a finite-design answer of a different kind. The useful open question is which identified Rasch functionals are decision-relevant and how a flexible posterior’s nonidentified directions affect their finite-sample interpretation.

23.3.2 A multivariate constrained Bayes and triple-goal

Multivariate use already exists. Shen and Louis state that their approaches generalize to multivariate unit-specific parameters, and Lockwood et al. apply CB and TG coordinatewise in a bivariate hierarchical model (Chapter 22). What is not supplied by those precedents is a canonical joint action: a transformation-aware loss, ordering or multivariate EDF target that is not an arbitrary collection of coordinatewise decisions. The open problem is to define and compare that joint target, not to invent multivariate CB/GR from nothing.

23.3.3 Misspecification of the measurement model

Every result in this book conditions on the item model being right. The whole programme asks what it costs to assume the wrong \(G\) while the Rasch or 2PL likelihood is correct.

Nothing here addresses the opposite failure, and there is reason to think it matters more. Xu et al.’s finding in the multidimensional case (Chapter 22) is that non-normality can distort which items load on which dimension — the misspecified population distribution damaging the measurement model, not just the recovered distribution. If that coupling has a unidimensional analogue, then the separation this book relies on — item side estimable without \(G\), so the cost of \(G\) is isolable — is a property of correctly specified item models and not of the estimator.

That would not invalidate anything proved here. It would relocate it.

23.4 A problem that is really a literature audit

Chapter 13 shows that a bimodal \(G\) can produce a unimodal sum-score distribution — total variation \(0.035\) on the worked case, with six spread items. The inference “the score distribution looks normal, therefore the trait distribution is” is therefore invalid, and it is common.

What is unknown is whether any published conclusion about latent non-normality rests on it. Answering that means reading a body of applied work with a specific question, which is a different kind of labour from the rest of this book and was not undertaken. It is listed because it is the item on this list most likely to change what someone does tomorrow.

23.5 The list

Table 23.1: What this book does not settle. Source: tables/T-open-problems.rds.
ID Problem Where it arose What is known Blocked on
OP-01 Finite-\(N\) behaviour of the DPM posterior for \(G\) under weak identification ch. 16 Under the anchored Rasch conditions the observable law is generated by the \(I+1\) displayed integral evaluations of \(G\), subject to their identities; guarantees for flexible priors are asymptotic The field
OP-02 Whether the rank proposition holds outside its stated conditions ch. 20 Reliability dominates and the shape residual is measurable; on the displayed grid it peaks near \(\bar w = 0.915\) and declines This book
OP-03 How far the incompatibility theorem generalizes ch. 17 An existence proof on two overlapping normals; the degenerate coincident case is exhibited This book
OP-04 A joint, transformation-aware multivariate CB/TG action ch. 18, 19, 22 Coordinatewise multivariate CB/TG exists; a canonical joint loss, ordering, or EDF action is not supplied The field
OP-05 What a finite polytomous test identifies about \(G\) ch. 22 The PCM keeps sufficiency and the GRM does not; the semiparametric question is not posed here This book
OP-06 How location/scale conventions transform priors on interpretable functionals of \(G\) ch. 4, 15, 16 Item centering and common atom translations leave \(K\) unchanged; other interpretable functionals can depend on location/scale convention This book
OP-07 What the normal working model costs at \(I < 15\) ch. 11 The exact posterior is computable; the working model is an approximation whose error is unquantified here This book
OP-08 Whether the sum-score fallacy changes a published conclusion ch. 13 Sum scores can be unimodal under a bimodal \(G\) (TV \(= 0.035\) on the worked case) A literature audit
OP-09 How much survives a misspecified measurement model ch. 3, 22 Every result here assumes the item model is correct The field

Five of the nine are marked as this book’s own. That count records responsibility, not difficulty. Some are bounded numerical checks; others require a new characterization or semiparametric identification argument whose difficulty has not been established. A reader deciding whether to trust Chapter 21’s contribution statement is entitled to know which questions the book scoped out, without being told that they are all tractable.

Three of the remaining four are genuinely open in the field, and one of those three — the finite-\(N\) meaning of a DPM posterior for \(G\) under weak identification — is the question that would do most to justify or undercut the entire flexible-prior programme this book describes.

23.6 Where the ledger leads

This chapter ends the Rasch-centred theory spine, not the volume. The next route depends on the reader’s question:

  • What changes when discriminations are free? Read Chapter 24 for information, Chapter 25 for reliability, and Chapter 26 for identification and the DPM. Together they are Part VIII, What Changes Under the 2PL.
  • What did the programme’s evidence settle? Read Chapter 27 for the simulation record, Chapter 28 for the real-test record, and Chapter 29 for the formal disposition of this ledger.