Tests are administered primarily to measure candidates' knowledge. How well do they master the material, and at what level? How did this group perform compared with other groups or with previous years? But it isn't only the candidates' knowledge that is measured, in effect, the test itself is being tested as well.
Personally, I find this one of the most interesting steps in the assessment cycle. Although there is a great deal you can do beforehand to build a valid and reliable test, only afterwards can you establish whether you succeeded. There are almost always a few questions that produce a different result than expected. Particularly when you have a sufficient number of administrations, it is worth examining both the individual questions and the test as a whole. Why did so many candidates answer a given question so poorly? Was the question itself incorrect? Was the wording ambiguous? Was it in fact testing different knowledge? Or was that knowledge not covered during teaching or not at the right level of application?
Is time taken to evaluate the test, identify points for improvement and actually implement them, so that the next group of candidates sits an even better test? Only then is the circle of the assessment cycle complete.
In a digital assessment system, statistical data becomes available for analysis almost as soon as the results come in, at both question and test level. These figures tell you something about the test and the underlying questions, but you always need to look at them more closely to understand what they mean in the context of that particular test.
At test level, you look at things like the mean and the spread of the results. Cronbach's alpha is then the best-known measure of a test's reliability: it measures internal consistency. A high score indicates that the questions in the test hang together and produce a reliable result; with a low score, it isn't really possible to determine whether the result says anything meaningful about how well candidates have mastered the material.
What you do need to watch out for is that Cronbach's alpha rises almost automatically as a test contains more questions, even if question quality has not improved. A high alpha can therefore also point to redundant (overly overlapping) questions rather than to genuine quality. The psychometric quality of the individual questions, their discriminating power and difficulty, also influences this reliability.
The validity of the test is equally essential: did the questions really assess the knowledge set out in the learning objectives, and at the right level? Did the test cover the material (content validity), and did each question actually measure the intended understanding rather than, say, language proficiency instead of subject knowledge (construct validity)? This also relates to constructive alignment, in which learning objectives, teaching activities and assessment are aligned with one another. It has to be judged on substance, separately from Cronbach's alpha, because a test can measure the wrong knowledge very consistently and still have a high Cronbach's alpha.
At question level, the p'-, Rit- and Rir-values give an indication of a question's quality, but a value that is (too) low or (too) high does not by definition mean the question is a poor one. It may well have been intentional to include the occasional (very) difficult question, in order to discriminate effectively between candidates who have and have not mastered the material, to assess this, you look at both the p'-value and the Rir-value. Before drawing conclusions from the statistics, however, you should always examine the question on its merits. Are the wording, the alternatives in a closed question and the assessment criteria in an open question all sound? Is there any indication of why candidates consistently chose one particular alternative? Is the question unambiguous? As with other steps in the assessment cycle, two pairs of eyes see more than one. Especially when you are analysing your own test and questions, it is valuable to have a colleague look with you, so you avoid blind spots. Alongside any adjustments to the marking scheme for the current candidates, this analysis also lets you improve the questions themselves for future administrations.
A test administration, then, is not finished once the results have been fed back to candidates. It offers an opportunity to work on the quality of future tests and with it, on the quality of teaching.
Do you take the time to evaluate and improve your tests?


.webp)

