Validity is the degree to which a test measures what it claims to measure. It is the most important quality of any assessment: a test that does not measure what it purports to is useless, regardless of how reliable or practical it is.
Modern testing theory (Messick 1989) treats validity as a unitary concept: there is not a set of separate "types" but rather different sources of evidence that support or undermine the overall validity argument. However, the traditional categories remain useful for thinking through test quality.
Does the test actually measure the theoretical construct it targets? A "reading comprehension" test that can be answered from general knowledge without reading the passage has weak construct validity. A speaking test where learners read aloud from a script does not measure spontaneous speaking ability.
Construct validity is the overarching concern; all other "types" contribute evidence for or against it. See Construct Validity for full treatment.
Does the test adequately sample the content domain it claims to cover? A listening test that only uses monologues and never dialogues under-represents the construct of listening ability. Content validity is established through expert judgment and specification matching, not statistics. See Content Validity for full treatment.
Does the test look like a legitimate test of what it claims to measure, to the people taking it? Face validity is technically not a "real" type of validity, but it matters pragmatically. A speaking test that feels like a speaking test motivates engagement; one that feels irrelevant breeds resentment and reduced effort. See Face Validity for full treatment.
Does test performance correlate with other measures of the same ability?
Reliability is a necessary but not sufficient condition for validity. A test can be perfectly reliable (consistent results every time) but completely invalid (consistently measuring the wrong thing). A grammar test given as a measure of speaking ability may be highly reliable but invalid for its stated purpose.
Conversely, a test cannot be valid if it is not reliable; inconsistent measurement cannot be accurate measurement.
| Threat | Example |
|---|---|
| Construct underrepresentation | A writing test that only assesses grammar, ignoring coherence, task achievement, and vocabulary |
| Construct-irrelevant variance | A reading test where scores depend heavily on background knowledge of the topic rather than reading ability |
| Method effects | Multiple-choice format allowing test-wise strategies unrelated to language ability |
| Bias | Test content that advantages certain cultural or gender groups |
| Washback misalignment | The test measures skills that do not align with the learning objectives |