Content validity is the extent to which a test's items and tasks adequately sample the content domain the test claims to cover. It asks: Does this test represent the full range of knowledge, skills, and abilities that define the domain?
Unlike construct validity, which relies on statistical evidence, content validity is established primarily through expert judgment: systematic evaluation of how well test content maps to the domain specification.
No test can assess everything within a domain. A "grammar test" cannot test every structure; a "reading test" cannot use every possible text type. Content validity is about whether the sample is representative of the population of possible items and tasks.
Hughes (2003) frames it as a specification-matching exercise:
| Aspect | Content validity | Construct Validity | Face Validity |
|---|---|---|---|
| Question | Does the test sample the domain adequately? | Does it measure the target ability? | Does it look right? |
| Method | Expert review against specifications | Statistical and theoretical analysis | Stakeholder impression |
| Judgments by | Subject matter experts, test developers | Researchers, psychometricians | Test-takers, administrators |
| Limitations | Cannot determine if test measures the right process | May miss domain sampling gaps | No technical basis |
Content validity contributes evidence toward construct validity but does not guarantee it. A test can sample content broadly (strong content validity) but still fail to measure the target ability; for instance, if all items test recognition rather than production, the content coverage is broad but the construct is narrowly operationalised.
A clear test specification (blueprint) defines:
Example for an end-of-course Reading and Listening test:
| Component | Proportion | Skills assessed |
|---|---|---|
| Listening: monologue | 25% | Gist, specific information, inference |
| Listening: dialogue | 25% | Specific information, attitude, opinion |
| Reading: long text | 30% | Main idea, detail, inference, vocabulary in context |
| Reading: short texts | 20% | Scanning, matching, specific information |
Independent experts evaluate items against the specification. Questions to ask (adapted from Hughes 2003):
Content validity is especially critical for achievement tests, because their purpose is to assess mastery of specific course content. An achievement test that omits key course objectives or over-represents minor topics has weak content validity, and produces invalid conclusions about learning.
The alignment chain should be:
Course objectives → Syllabus content → Test specification → Test items
Breakdowns at any point weaken content validity. Common failures:
For proficiency tests, the content domain is defined by a theory of language ability rather than a syllabus. IELTS, for example, must represent the domain of "academic English ability" through its choice of text types, task types, topics, and scoring criteria. Content validity questions include:
Content validity is the most accessible form of validity evidence for classroom teachers. You do not need statistics; you need a clear specification and the willingness to check your test against it.
Practical implications: