Humans-in-the-Loop

As AIME-Con approaches, my thoughts turn to automated item generation (AIG) and plans to reduce the many human roles to a human-in-the-loop role.

Obviously, psychometric reviews are easily automatable. I mean, you wouldn’t call for automating other work you’re not actually expert in if you didn’t think all the field test data processing, all the psychometric forms reviews and all the final score determinations were already easily automatable, right? 

So, we can check off those three steps of psychometric work as automatable. 

What other points or types of review are automatable? You looked into all the issues of graphics review, right? You made sure that that was all automatable with LLM-based tools?

Blueprint conformity is pretty easy, right? Representation across the content, all the different types of items and stimulus types called for in the blueprint? That’s all in the metadata and you just need to match that up to blueprint requirements, right?

You do the accessibility reviews? Which ones did you check? You sure they are all automatable?

Fairness reviews? Sensitivity? Bias? Representation? You have a good balance of political/cultural sensitivity and authenticity, right? Just curious, which cultural perspectives do your tools take into account? 

ELA’s stimulus reviews? You’ve got your various fairness reviews? Content reviews? You’re satisfied that LLM-based technologies can identify grade-level testable points? 

How about client reviews? You know, before stimulus reviews, before item external reviews, before field testing, before operational testing. Are those going to be automated?

What about reviews of modified material for test takers with sensory and motor disabilities? Oh, I don’t mean the earlier accessibility reviews. I mean for the modified stuff. You accounting for that? 

Inter-item clueing? 

I’m curious about the teacher reviews. You know, the ones that enable the industry, state officials and politicians to say things like “Every question on our tests was reviewed by a team of teachers from our state.” You automating that? You’re automating the easier item starter/drafting stages, right?  The ones that allow state officials and politicians to (misleadingly) say, “All the items on our tests were written by teachers from our state.” You automating that? 

What about reviewing items for cognitive complexity and whether they are hitting the core vs. the margins of each alignment reference? Is your system reviewing for that? 

OK. Just making sure your plan for a human in the loop is accounting for all the human reviews and all the types of expertise that are so important to ensuring that we have items that elicit evidence of the targeted cognition for the range of typical test takers.