Construct validity is the extent to which a test actually measures the theoretical ability (construct) it claims to measure. It is the central, unifying concept in modern validity theory; Messick (1989) argued that all validity is ultimately construct validity, with other "types" (content, criterion-related) providing different sources of evidence for the construct validity argument.
Does a test score mean what we say it means?
If a test is labelled "reading comprehension," construct validity asks: Do the scores genuinely reflect reading comprehension ability, or do they reflect something else: background knowledge, test-taking strategy, vocabulary size, or working memory capacity?
Messick (1989) identified two threats that undermine construct validity:
The test is too narrow; it fails to capture important aspects of the construct. Examples:
The test scores are influenced by factors outside the target construct. Two subtypes:
| Aspect | Construct validity | Content Validity | Face Validity |
|---|---|---|---|
| Focus | Does it measure the right ability? | Does it sample the domain adequately? | Does it look right? |
| Evidence | Statistical, theoretical, logical | Expert judgment, specification matching | Stakeholder impression |
| Status | The overarching validity concept | One source of evidence for construct validity | Not technically validity |
| Example | Factor analysis shows speaking scores load on one dimension, not separate grammar/fluency factors | The test covers all four language skills proportionally | Test-takers feel the speaking test is a fair test of speaking |
Content validity and face validity are contributory evidence for construct validity, not separate concepts at the same level. A test with strong content validity (good domain sampling) and strong face validity (stakeholder acceptance) has some evidence supporting construct validity, but not conclusive evidence.
Correlation studies. If a test measures reading comprehension, scores should correlate with other established measures of reading comprehension (convergent validity) and correlate less strongly with measures of unrelated abilities (discriminant validity). The classic framework is Campbell & Fiske's (1959) multitrait-multimethod matrix.
Factor analysis. Statistical analysis of score patterns reveals whether the test measures one ability or several. If a "reading test" turns out to measure vocabulary knowledge and reading speed as separate factors, this informs the construct definition.
Differential item functioning (DIF). Analysis of whether items function differently for subgroups (gender, L1, cultural background). If an item is significantly harder for one group than another at the same ability level, construct-irrelevant variance is present.
Think-aloud protocols. Observing what test-takers actually do when answering items: are they using the target skill or something else? If "reading comprehension" items can be answered through vocabulary matching without actual comprehension, the construct is undermined.
Expert judgment. Subject matter experts evaluate whether test tasks require the target ability. This overlaps with content validation but focuses specifically on the cognitive processes demanded.
Construct validity is not an abstract concern; it has direct consequences for teaching and learning:
For test developers and teachers alike, the question is always: Am I measuring what I think I am measuring, or am I measuring something else?