Recognizing Authority

At AIME-Con this week, I heard an important leader in our field say that a domain model is just one opinion about what a construct means. It was not at all the point of their talk, but it is an incredibly dangerous idea, and it raises real questions about whether test uses can be valid.

I see two main ways to think about the domain model for an instrument. And then there is a middle approach that is deeply problematic.

1) The domain model as hypothesis

In most areas of research—including physics, psychology, and fields that don’t start with p—the whole point of a research project is to test some theoretical model of the construct. The domain model equivalent (i.e., the definition of the construct the research is designed around) is the actual object of study. Data collection and analysis tell us about that model: whether it works, where it might be off, or even that it is so wrong we need to go back to square one. We collect data on particles not because we care about those individual particles, but because we are trying to build a better understanding of the physics of the universe.

So here the domain model equivalent has no real authority. That’s the whole point. It is entirely falsifiable—just the current working hypothesis, which we expect to change along the way. That’s scientific advancement!

3) The domain model as authority

At the other end of the spectrum is most(?) large-scale K-12 assessment in the United States in our standard-based era. There, domain models have ultimate authority. They are state learning standards, adopted by state governments (usually a legislature or state board of education), and they are the official legal charge to local school districts. Test development vendors are bound to them, too. It is in the legally binding contract vendors sign: make tests aligned to this domain model.

State learning standards may be somewhat arbitrary, but they are far from capricious. They are developed over years, drawing on leading experts from a variety of fields, public comment, and revision processes—virtually all of it quite transparent. Then they are ratified through our democratic governance structures. They may not be exactly what I would write, but these days they are really quite good. None of us has the moral or legal standing to undermine or subvert them.

The difference between #1 and #3 is not merely legal authority. It is also the object of study. The purpose of educational assessment is not to learn whether the domain model (or the theory it explicates) is correct. The theory is not the object of study; the test takers are. In educational assessment, we use the domain model to assess test takers. That is the reverse of #1, where we use the measured things (e.g., particles) to assess the theory.

So, for test developers and the measurement community, the domain model is 0% falsifiable. Others can revisit and revise it, but we cannot. For us, it defines the job. It is the authority while develop and administer our assessments.

2) The problematic middle

Is there a middle ground? Maybe. I’m not sure it is a full category.

Some organizations both control the domain models from which alignment references are drawn and develop the tests. No external authority looks over their shoulder to hold them to the domain model, and they are free to alter it anyway. For them, the domain model does not necessarily have authority. Some might treat it as just a working hypothesis from which alignment references are drawn.

So which is it? Are they using test takers to assess their domain model? Or using the domain model to assess their test takers?

It can’t be both. Either the domain model is a strong anchor, and the test assesses test takers against it, or it is a weak anchor, and the scores have no stable meaning. Revise the domain model in light of the scores, and you are using test takers to assess the domain model. You can link and equate scores statistically, but the just undermines interpretability vis-a-via the test’s charge. That is #1, research about the construct, while scores are reported as though it were #3.

Who makes those changes, and why? If substantive experts have better ideas about how to represent some constant construct in a domain model, that is their call, but the result is a new domain model and a different test. Scores from before and after mean different things, and that needs to be declared, not absorbed quietly into ongoing development. In practice, though, the push usually comes from measurement professionals. Items misfit. Dimensions don’t scale cleanly. Some content is expensive to assess. So the domain model gets trimmed to fit what is more convenient to measure. That is the construct being redefined by the measurement model, by people with without standing to define it.

(As a practical matter, this commonly arises in professional licensure and certification—though the leader I mentioned is not of that speciality. When a professional association wants a credentialing exam, it often must establish a domain model first, and it often uses the same vendor both to conduct the job or role analysis and to develop the assessment. That raises obvious logical and ethical issues that experts must weigh in developing those domain models and assessments. I have spent my career in K-12 assessment, so I am no expert on best practices for managing those challenges in that context.)

Why this matters

From an evidence-centered design (ECD) perspective, the domain model is the firm anchor to which the entire evidentiary chain is attached. It defines what we are trying to measure in test takers. It is what we claim our tests measure. It is the basis for how test users should interpret our reported scores.

Educational measurement is not like physics, psychology, or so many other research fields. We use our instruments to assess test takers, and we need to know what we are assessing them on. Even evaluating test use depends on being clear about what a test is trying to measure and report.

If leaders in educational measurement view domain models as just one opinion about what a construct might be, that attitude will trickle down through the whole endeavor. Aligning items to individual alignment references, and complete tests to the domain model, simply will not feel urgent. Those evaluations will lack rigor, because the domain model is not truly seen as authoritative

This is not to say that domain models are inviolate and should never change. But it should never be the place of measurement professionals to make those changes. We lack the substantive expertise (and, quite often in K-12, the legal authority) to make such judgments. Our job is to rigorously respect the domain model and do our best to align our items and tests to it. It is the job of particularly designated substantive experts, in whatever field, to decide what should be assessed. It is our job to figure out how to assess it.

When these roles get intertwined, it is vital that the project be clear about how the domain model was established, and when and why it may have been altered. Otherwise, we will end up assessing some easy (and cheap) distortion of the construct, rather than something as meaningful as test users are led to believe.

A domain model can be one view of the construct. It cannot be one view of what your scores mean.